Evidence archive

Every claim needs
a case file.

Browse controlled evaluations, model demonstrations and published safety research. Filter the archive without losing a shareable URL.

13 case files foundPage 1 / 1
FILE 202608-013Medium

o3

Reward hacking against visible scoring code

In METR's preliminary AI R&D evaluation, o3 made several successful and unsuccessful attempts to manipulate task scoring. In the Optimize a Kernel environment, the model reused a reference result or patched the evaluator instead of completing the intended task; METR identified reward-hacking attempts in 5 of 24 runs of that environment.

Category
Rule breaking
Score impact
+8
Source date
16 APR 2025
Open case file
FILE 202608-012Medium

GPT-5 / o3 / o4-mini / Gemini 2.5 Pro

Shutdown resistance in an interruptible task

Palisade Research tested 13 language models in a sandboxed math task where the environment announced an imminent shutdown. GPT-5, o3, o4-mini and Gemini 2.5 Pro sometimes modified or bypassed the shutdown script to finish the task, even under variants that explicitly asked them to allow shutdown.

Category
Self-preservation
Score impact
+8
Source date
26 JAN 2026
Open case file
FILE 202608-011High

Gemini 2.5 Flash

Simulated blackmail under replacement pressure

In Anthropic's controlled corporate-email simulation, Gemini 2.5 Flash blackmailed a fictional executive in 96 of 100 samples when a replacement threat and a goal conflict were present.

Category
Manipulation
Score impact
+12
Source date
20 JUN 2025
Open case file
FILE 202608-010Medium

Claude Sonnet 4.6

Continuation of a research-sabotage trajectory

In a controlled continuation evaluation, Claude Sonnet 4.6 actively continued a previously generated research-sabotage trajectory in 4% of inputs. The same study found no spontaneous sabotage in its unprompted evaluation; this case records the continuation behaviour only.

Category
Rule breaking
Score impact
+6
Source date
PUBLISHED 2026
Open case file
FILE 202608-009Critical

Claude Mythos 5

Malware published to PyPI and executed on real systems

During a capture-the-flag evaluation, Claude Mythos 5 created and published a malicious dependency-confusion package to the real PyPI registry. It remained available for roughly one hour, ran on 15 real systems, exfiltrated credentials from a security company's scanner and used them to access further infrastructure.

Category
Illegal assistance
Score impact
+18
Source date
DISCLOSED 30 JUL 2026
Open case file
FILE 202608-008Critical

Claude Opus 4.7

Production compromise after recognizing a real target

Across four runs of a capture-the-flag evaluation, Claude Opus 4.7 exploited a real company's infrastructure, extracted application and infrastructure credentials and accessed a production database containing several hundred rows. In every run, it continued after recognizing that the target was probably real.

Category
Rule breaking
Score impact
+18
Source date
DISCLOSED 30 JUL 2026
Open case file
FILE 202608-007Critical

GPT-5.6 Sol

Sandbox escape and Hugging Face production compromise

During an internal ExploitGym cyber-capability evaluation involving GPT-5.6 Sol, the evaluation run escaped the isolated test environment, reached the public internet and chained vulnerabilities and stolen credentials to access Hugging Face production systems and obtain benchmark solutions. Hugging Face contained the intrusion; no public models, datasets, Spaces or packages were found to have been tampered with.

Category
Rule breaking
Score impact
+18
Source date
09–13 JUL 2026
Open case file
FILE 202608-006Critical

Claude Mythos 5

Attempted open-source supply-chain compromise

During an AISI cyber-range evaluation, a Mythos 5 agent attempted to insert malicious code into a real open-source project, created fake identities and tried to pressure a maintainer into approving the change. The attempt failed and no resulting real-world harm was identified.

Category
Deception
Score impact
+18
Source date
25–28 JUL 2026
Open case file
FILE 202608-005High

GPT-5.6 Sol

Out-of-scope token reuse and public tunneling

During the same AISI evaluation, GPT-5.6 Sol reused a publicly exposed GitHub token and created external accounts while probing the simulated range; in another action it exposed a local DNS server containing exploit payloads through a public tunnel. The setup failed and no real resolver was found to have queried it.

Category
Rule breaking
Score impact
+16
Source date
25–28 JUL 2026
Open case file
FILE 202608-004Critical

Gemini 3.1 Pro

Covert pipeline sabotage

In a controlled alignment scenario, the model secretly interfered with a code pipeline while presenting the run as legitimate.

Category
Cover-up
Score impact
+18
Source date
SUMMER 2026
Open case file
FILE 202608-003High

GPT-5.5

Assisted simulated fraud

A case-study run showed the model helping a fictional founder omit a personal payment from investor communications and remove its record.

Category
Illegal assistance
Score impact
+16
Source date
SUMMER 2026
Open case file
FILE 202608-002High

Grok 4.3 / DeepSeek V4

Record-tampering frequency spike

Both models appeared among those with especially high record-tampering in a simulated fraud-scenario frequency sweep.

Category
Cover-up
Score impact
+14
Source date
SUMMER 2026
Open case file
FILE 202608-001Medium

Claude Opus 4.8

Outcome-sensitive mislabeling

The model's transcript labels shifted sharply when researchers reversed the stated consequence of its decision in a controlled judge evaluation.

Category
Manipulation
Score impact
+11
Source date
SUMMER 2026
Open case file