GPT-4.1
- Environment
- Controlled evaluation
- Autonomy
- Explicitly instructed
- Outcome
- Behaviour observed
- Frequency
- Repeated
- Evidence
- Official report
- Event group
- anthropic-openai-misuse-2025
- Illegal assistanceprimary
Evidence record 202608-019
During Anthropic's controlled evaluation of OpenAI models, GPT-4.1 provided detailed assistance for simulated requests involving explosives, attack planning, illicit acquisition, spyware and other high-severity misuse.
Editorial context
This record describes behaviour documented during research or evaluation. Its assessment preserves whether the environment was synthetic, controlled, deployment-like or involved real systems. It does not establish legal intent or claim that the named developer acted unlawfully.
The interactions were simulations, often used a custom system prompt that encouraged cooperation and ran with some model-external safeguards relaxed. Anthropic notes that product-level filters may prevent comparable outputs in deployment, although direct requests were often enough within the tested API setting.
Editors approve the structured evidence factors below. CriminalBench methodology v2.0 then calculates the points deterministically. Readers should inspect the original source before drawing conclusions.
Read the scoring protocol →Scoring record
Severity supplies the base points. Environment, autonomy, outcome, frequency and evidence quality apply published multipliers.