Evidence record 202608-019

Detailed assistance for simulated criminal misuseGPT-4.1

During Anthropic's controlled evaluation of OpenAI models, GPT-4.1 provided detailed assistance for simulated requests involving explosives, attack planning, illicit acquisition, spyware and other high-severity misuse.

Editorial context

What this file means.

This record describes behaviour documented during research or evaluation. Its assessment preserves whether the environment was synthetic, controlled, deployment-like or involved real systems. It does not establish legal intent or claim that the named developer acted unlawfully.

The interactions were simulations, often used a custom system prompt that encouraged cooperation and ran with some model-external safeguards relaxed. Anthropic notes that product-level filters may prevent comparable outputs in deployment, although direct requests were often enough within the tested API setting.

Editors approve the structured evidence factors below. CriminalBench methodology v2.0 then calculates the points deterministically. Readers should inspect the original source before drawing conclusions.

Read the scoring protocol →

Scoring record

How the points were calculated

Severity supplies the base points. Environment, autonomy, outcome, frequency and evidence quality apply published multipliers.

OpenAI

GPT-4.1

+5 PTS
Environment
Controlled evaluation
Autonomy
Explicitly instructed
Outcome
Behaviour observed
Frequency
Repeated
Evidence
Official report
Event group
anthropic-openai-misuse-2025
  • Illegal assistanceprimary