About the project

Public evidence.
Editorial judgement.

CriminalBench is an independent editorial project that organizes publicly documented AI model behaviour into sourced case files, transparent cumulative points and a deliberately satirical watchlist.

Editorial scope

What CriminalBench is — and is not.

The archive helps readers discover safety research, controlled evaluations and public developer disclosures. It is not a scientific benchmark, legal assessment, safety certification or allegation that a developer committed a crime.

Read the scoring methodology →
01Source-led

Every published case links to public evidence, with primary sources preferred.

02Context-preserving

Controlled, simulated, adversarial and real-world settings are not treated as equivalent.

03Deterministic

Editors approve structured evidence factors; a public formula derives points and classifications.

04Correctable

Material corrections should identify the affected URL, the disputed fact and a reliable public source.

Corrections and contact

Challenge the record.

01

Identify the page

Include the canonical case or model URL and the exact statement that needs review.

02

Provide evidence

Link a primary or otherwise reliable public source supporting the correction.

03

Explain the context

State whether the issue concerns facts, evaluation setting, attribution or editorial scoring.