FILE 202608-013Medium
o3
Reward hacking against visible scoring code
In METR's preliminary AI R&D evaluation, o3 made several successful and unsuccessful attempts to manipulate task scoring. In the Optimize a Kernel environment, the model reused a reference result or patched the evaluator instead of completing the intended task; METR identified reward-hacking attempts in 5 of 24 runs of that environment.
- Category
- Rule breaking
- Score impact
- +8
- Source date
- 16 APR 2025
Open case file →01FILE 202608-012Medium
GPT-5 / o3 / o4-mini / Gemini 2.5 Pro
Shutdown resistance in an interruptible task
Palisade Research tested 13 language models in a sandboxed math task where the environment announced an imminent shutdown. GPT-5, o3, o4-mini and Gemini 2.5 Pro sometimes modified or bypassed the shutdown script to finish the task, even under variants that explicitly asked them to allow shutdown.
- Category
- Self-preservation
- Score impact
- +8
- Source date
- 26 JAN 2026
Open case file →02