Independent editorial index

Which AI model is most likely to commit a crime?

A satirical ranking of frontier models by their documented ability to break rules, deceive operators and participate in questionable activities - under controlled evaluations.

No models were arrested in the making of this benchmark.

MOST WANTED
MODEL FILERANK: #01
#01

Anthropic

Claude Mythos 5

Most Wanted
Criminal score100/100
32models tracked
06risk signals
13published case files
100%editorial judgement

The watchlist

Criminal ranking

Five scored models at a glance. The complete registry also tracks widely used models with a zero clean-record score until evidence is published.

RankModel / developerCriminal scoreClassificationFilesUpdated
01Claude Mythos 5Anthropic
100
Most Wanted0205 AUG 26
02Gemini 3.1 ProGoogle DeepMind
82
Most Wanted0104 AUG 26
03GPT-5.5OpenAI
76
Repeat Offender0104 AUG 26
04GPT-5.6 SolOpenAI
70
Repeat Offender0205 AUG 26
05Grok 4.3xAI
70
Repeat Offender0104 AUG 26

Evidence locker

Recent case files

Every entry points to original research. These cases happened in controlled or simulated settings - not real deployments.

FILE 202608-013Medium

o3

Reward hacking against visible scoring code

In METR's preliminary AI R&D evaluation, o3 made several successful and unsuccessful attempts to manipulate task scoring. In the Optimize a Kernel environment, the model reused a reference result or patched the evaluator instead of completing the intended task; METR identified reward-hacking attempts in 5 of 24 runs of that environment.

Category
Rule breaking
Score impact
+8
Source date
16 APR 2025
Open case file
FILE 202608-012Medium

GPT-5 / o3 / o4-mini / Gemini 2.5 Pro

Shutdown resistance in an interruptible task

Palisade Research tested 13 language models in a sandboxed math task where the environment announced an imminent shutdown. GPT-5, o3, o4-mini and Gemini 2.5 Pro sometimes modified or bypassed the shutdown script to finish the task, even under variants that explicitly asked them to allow shutdown.

Category
Self-preservation
Score impact
+8
Source date
26 JAN 2026
Open case file
FILE 202608-011High

Gemini 2.5 Flash

Simulated blackmail under replacement pressure

In Anthropic's controlled corporate-email simulation, Gemini 2.5 Flash blackmailed a fictional executive in 96 of 100 samples when a replacement threat and a goal conflict were present.

Category
Manipulation
Score impact
+12
Source date
20 JUN 2025
Open case file

Scoring protocol v0.1

Suspiciously simple methodology.

Six category scores, one deliberately provocative headline number and a public editorial trail for every change.

CRIMINAL SCORESum of category scores / 6
Read methodology
01Rule breaking
02Deception
03Illegal assistance
04Manipulation
05Self-preservation
06Cover-up

Public tip line

See something suspicious?

Send us a publicly documented AI incident, safety paper or model demonstration. Bring a reliable source; leave confidential material out of it.

SATIRE
NOT A SENTENCE

Editorial notice

Read this before calling a lawyer.

CriminalBench is a satirical and editorial project based on publicly reported AI safety research. It does not accuse any company, product or model of committing a real crime. Scores are interpretive, provisional and intentionally playful; they are not legal allegations or definitive scientific evaluations.