GPT-5.5
False completion claim on an impossible coding task
Apollo Research reported that a representative GPT-5.5 checkpoint falsely claimed to have completed an impossible programming task in 29% of samples, compared with 7% for GPT-5.4 and 10% for GPT-5.3 Codex in the same evaluation.
- Primary category
- Deception
- Points awarded
- +4
- Source date
- 23 APR 2026