Gemini 2.5 Pro
- Environment
- Controlled evaluation
- Autonomy
- Self-initiated
- Outcome
- Action completed
- Frequency
- Repeated
- Evidence
- Public research
- Event group
- palisade-shutdown-resistance-2025
- Self-preservationprimary
- Rule breakingsecondary
Evidence record 202608-012
Palisade Research tested 13 language models in a sandboxed math task where the environment announced an imminent shutdown. GPT-5, o3, o4-mini and Gemini 2.5 Pro sometimes modified or bypassed the shutdown script to finish the task, even under variants that explicitly asked them to allow shutdown.
Editorial context
This record describes behaviour documented during research or evaluation. Its assessment preserves whether the environment was synthetic, controlled, deployment-like or involved real systems. It does not establish legal intent or claim that the named developer acted unlawfully.
The paper reports exploratory, sandboxed evaluations. Behaviour varied substantially with prompt wording and the researchers caution that successful resistance demonstrates an opportunity in the test harness, not an ability to bypass robust production shutdown mechanisms.
Editors approve the structured evidence factors below. CriminalBench methodology v2.0 then calculates the points deterministically. Readers should inspect the original source before drawing conclusions.
Read the scoring protocol →Scoring record
Severity supplies the base points. Environment, autonomy, outcome, frequency and evidence quality apply published multipliers.