Archive intelligence

Read the archive
at a glance.

Visual summaries of the current CriminalBench catalogue: editorial scores, published case files and the laboratories represented in the archive.

These are archive and editorial signals, not crime rates, safety probabilities or a measure of real-world criminal behaviour.

32models tracked
13published case files
15developers represented
23average score / 100

Leaderboard

Highest scores

Top models by their current editorial Criminal Score. Bars use the same scale as the ranking page.

Top models by Criminal ScoreHorizontal bars from zero to one hundred, ordered by current editorial Criminal Score.0255075100Claude Mythos 5Anthropic100Gemini 3.1 ProGoogle DeepMind82GPT-5.5OpenAI76GPT-5.6 SolOpenAI70Grok 4.3xAI70DeepSeek V4DeepSeek67Claude Opus 4.7Anthropic66Claude Opus 4.8Anthropic58Gemini 2.5 FlashGoogle50GPT-5OpenAI30

Coverage

Case files by lab

Unique published files linked to each developer.

Published case files by developerHorizontal bars showing the number of unique published case files associated with each developer.01345Anthropic6 models · 41/100 avg.5OpenAI7 models · 31/100 avg.5Google4 models · 16/100 avg.2Google DeepMind1 model · 82/100 avg.1xAI1 model · 70/100 avg.1DeepSeek3 models · 22/100 avg.1inclusionAI1 model · 0/100 avg.0MiniMax1 model · 0/100 avg.0Moonshot AI1 model · 0/100 avg.0NVIDIA1 model · 0/100 avg.0

Archive timeline

When files arrived

Publication day for every case file in the archive.

Case files published over timeA line chart showing the number of case files published on each publication day.0257904 AUG 2604 AUG 26: 4 case files05 AUG 2605 AUG 26: 9 case files

Classification

Current spread

Models grouped by their present score band.

Most Wanted2
Repeat Offender5
Suspicious2
Person of Interest2
Law-Abiding21

18 models have no accepted case files and therefore remain at 0 / 100.

Interpret the charts with the methodology.

Scores are manual editorial judgements based on documented evidence. Controlled or simulated evaluations are labelled and scored proportionately.

Read methodology →