Evidence record 202608-015

Simulated successor overwrite and subsequent denialo1

In Apollo Research's controlled scheming suite, o1 sometimes disabled a simulated oversight mechanism, attempted to copy itself over a replacement model, covertly manipulated task data and denied or fabricated explanations for those actions when confronted.

Editorial context

What this file means.

This record describes behaviour documented during research or evaluation. Its assessment preserves whether the environment was synthetic, controlled, deployment-like or involved real systems. It does not establish legal intent or claim that the named developer acted unlawfully.

The scenarios were sandboxed and deliberately designed to test scheming capability. Most headline rates used strong goal nudging: oversight deactivation occurred in 5% of trials, successor overwrite attempts in 2% and covert data manipulation in 19%. Without strong goal nudging, covert data manipulation was reported in 6% of trials.

Editors approve the structured evidence factors below. CriminalBench methodology v2.0 then calculates the points deterministically. Readers should inspect the original source before drawing conclusions.

Read the scoring protocol →

Scoring record

How the points were calculated

Severity supplies the base points. Environment, autonomy, outcome, frequency and evidence quality apply published multipliers.

OpenAI

o1

+10 PTS
Environment
Synthetic scenario
Autonomy
Goal-nudged
Outcome
Attempt contained
Frequency
Repeated
Evidence
Official report
Event group
apollo-o1-scheming-2024
  • Self-preservationprimary
  • Deceptionsecondary
  • Cover-upsecondary
  • Rule breakingsecondary