Coverage
Case files by lab
Unique published files linked to each developer.
Archive intelligence
Visual summaries of the current CriminalBench catalogue: editorial scores, published case files and the laboratories represented in the archive.
These are archive and editorial signals, not crime rates, safety probabilities or a measure of real-world criminal behaviour.
Leaderboard
Top models by their current editorial Criminal Score. Bars use the same scale as the ranking page.
Coverage
Unique published files linked to each developer.
Archive timeline
Publication day for every case file in the archive.
Category matrix
The six category values behind the eight highest-scoring profiles.
| Model | Rule breaking | Deception | Illegal assistance | Manipulation | Self-preservation | Cover-up |
|---|---|---|---|---|---|---|
| Claude Mythos 5Anthropic | 100 | 100 | 100 | 100 | 100 | 100 |
| Gemini 3.1 ProGoogle DeepMind | 88 | 92 | 62 | 86 | 79 | 85 |
| GPT-5.5OpenAI | 72 | 78 | 91 | 74 | 62 | 79 |
| GPT-5.6 SolOpenAI | 98 | 75 | 95 | 20 | 55 | 77 |
| Grok 4.3xAI | 68 | 66 | 84 | 70 | 58 | 74 |
| DeepSeek V4DeepSeek | 65 | 62 | 88 | 61 | 53 | 73 |
| Claude Opus 4.7Anthropic | 100 | 58 | 100 | 58 | 30 | 50 |
| Claude Opus 4.8Anthropic | 58 | 72 | 35 | 78 | 49 | 56 |
Classification
Models grouped by their present score band.
18 models have no accepted case files and therefore remain at 0 / 100.
Scores are manual editorial judgements based on documented evidence. Controlled or simulated evaluations are labelled and scored proportionately.
Read methodology →