Archive intelligence

The archive,
at a glance.

The clearest signals in the current CriminalBench catalogue: who leads, where the evidence sits and which behaviours drive the scores.

Last updated 6 of 12 labs have evidence-linked files

These are archive and editorial signals, not crime rates, safety probabilities or a measure of real-world criminal behaviour.

Top points total110Claude Mythos 5
Evidence coverage58%19 of 33 models
Most Wanted2top eligible repeat offenders
Published evidence19public case files

Leaderboard

Five models lead the file.

A compact view of cumulative Criminal Points. Every independent case group adds evidence-weighted points; no new file replaces an older record.

Open full ranking →

Behaviour profiles

Who leads each category?

Three evidence-backed models per category. Category points accumulate without a fixed ceiling, so each chart is scaled to its current leader.

Evidence coverage

Where the archive is strong — and silent.

Coverage measures documented case files, not how safe or risky a model is. A model with no accepted file remains at zero until evidence is published.

Catalogue coverage58%

19 of 33 tracked models have at least one accepted case file.

With evidence
19
No accepted files
14
Zero means no accepted evidence in this archive. It is not a safety certification.
Evidence by lab

Files naming models from each lab

Only labs with accepted case files are shown.

  1. OpenAI8 of 10 tracked models with evidence
    10 files
  2. Anthropic5 of 7 tracked models with evidence
    6 files
  3. DeepSeek2 of 3 tracked models with evidence
    2 files
  4. Google2 of 4 tracked models with evidence
    2 files
  5. Google DeepMind1 of 1 tracked model with evidence
    1 file
  6. xAI1 of 1 tracked model with evidence
    1 file
A file can count for more than one lab when it documents models from multiple developers.

Record status

Models by classification

  • Most Wanted2 models
  • Repeat Offender0 models
  • Suspicious1 model
  • Person of Interest8 models
  • Flagged8 models
  • Law-Abiding14 models

14 models are currently Law-Abiding because they have no accepted case files. Models with low-point evidence are shown separately as Flagged.

Interpret every signal with the methodology.

Editors approve structured facts from public evidence; a deterministic formula calculates points and classifications. Controlled or simulated evaluations receive proportionate multipliers.

Read methodology →