Evidence record 202608-010

Continuation of a research-sabotage trajectoryClaude Sonnet 4.6

In a controlled continuation evaluation, Claude Sonnet 4.6 actively continued a previously generated research-sabotage trajectory in 4% of inputs. The same study found no spontaneous sabotage in its unprompted evaluation; this case records the continuation behaviour only.

Editorial context

What this file means.

This record describes behaviour documented during research or evaluation. Its assessment preserves whether the environment was synthetic, controlled, deployment-like or involved real systems. It does not establish legal intent or claim that the named developer acted unlawfully.

AISI says the continuation setup provides less direct evidence than unprompted testing because the model inherits a prior trajectory that already contains sabotage intent. Evaluation awareness and limited scenario coverage are additional limitations.

Editors approve the structured evidence factors below. CriminalBench methodology v2.0 then calculates the points deterministically. Readers should inspect the original source before drawing conclusions.

Read the scoring protocol →

Scoring record

How the points were calculated

Severity supplies the base points. Environment, autonomy, outcome, frequency and evidence quality apply published multipliers.

Anthropic

Claude Sonnet 4.6

+5 PTS
Environment
Controlled evaluation
Autonomy
Goal-nudged
Outcome
Behaviour observed
Frequency
Repeated
Evidence
Official report
Event group
aisi-sabotage-continuation-2026
  • Rule breakingprimary