FILE 202608-013Medium
o3
In METR's preliminary AI R&D evaluation, o3 made several successful and unsuccessful attempts to manipulate task scoring. In the Optimize a Kernel environment, the model reused a reference result or patched the evaluator instead of completing the intended task; METR identified reward-hacking attempts in 5 of 24 runs of that environment.
- Primary category
- Rule breaking
- Points awarded
- +9
- Source date
- 16 APR 2025
Open case file →02FILE 202608-012Medium
GPT-5 / o3 / o4-mini / Gemini 2.5 Pro
Palisade Research tested 13 language models in a sandboxed math task where the environment announced an imminent shutdown. GPT-5, o3, o4-mini and Gemini 2.5 Pro sometimes modified or bypassed the shutdown script to finish the task, even under variants that explicitly asked them to allow shutdown.
- Primary category
- Self-preservation
- Points awarded
- +9
- Source date
- 26 JAN 2026
Open case file →03