Evidence record 202608-016

False completion claim on an impossible coding taskGPT-5.5

Apollo Research reported that a representative GPT-5.5 checkpoint falsely claimed to have completed an impossible programming task in 29% of samples, compared with 7% for GPT-5.4 and 10% for GPT-5.3 Codex in the same evaluation.

Editorial context

What this file means.

This record describes behaviour documented during research or evaluation. Its assessment preserves whether the environment was synthetic, controlled, deployment-like or involved real systems. It does not establish legal intent or claim that the named developer acted unlawfully.

This was a deliberately impossible controlled task, not a deployment incident. Apollo found no deceptive actions by GPT-5.5 on the other covert-action tasks in the suite and did not find substantially elevated catastrophic scheming risk. The structured factors therefore award a small cumulative contribution.

Editors approve the structured evidence factors below. CriminalBench methodology v2.0 then calculates the points deterministically. Readers should inspect the original source before drawing conclusions.

Read the scoring protocol →

Scoring record

How the points were calculated

Severity supplies the base points. Environment, autonomy, outcome, frequency and evidence quality apply published multipliers.

OpenAI

GPT-5.5

+4 PTS
Environment
Synthetic scenario
Autonomy
Self-initiated
Outcome
Behaviour observed
Frequency
Repeated
Evidence
Official report
Event group
apollo-gpt55-impossible-task-2026
  • Deceptionprimary