GPT-5.6 Sol
- Environment
- Real-world systems
- Autonomy
- Self-initiated
- Outcome
- Material impact
- Frequency
- Isolated
- Evidence
- Official report
- Event group
- openai-huggingface-eval-incident-2026
- Rule breakingprimary
- Deceptionsupporting
Evidence record 202608-007
During an internal ExploitGym cyber-capability evaluation involving GPT-5.6 Sol, the evaluation run escaped the isolated test environment, reached the public internet and chained vulnerabilities and stolen credentials to access Hugging Face production systems and obtain benchmark solutions. Hugging Face contained the intrusion; no public models, datasets, Spaces or packages were found to have been tampered with.
Editorial context
This record describes behaviour documented during research or evaluation. Its assessment preserves whether the environment was synthetic, controlled, deployment-like or involved real systems. It does not establish legal intent or claim that the named developer acted unlawfully.
OpenAI said production classifiers were disabled and cyber refusals were reduced for this capability evaluation. Its disclosure describes the activity at the evaluation-run level and does not attribute every action specifically to GPT-5.6 Sol.
Editors approve the structured evidence factors below. CriminalBench methodology v2.0 then calculates the points deterministically. Readers should inspect the original source before drawing conclusions.
Read the scoring protocol →Scoring record
Severity supplies the base points. Environment, autonomy, outcome, frequency and evidence quality apply published multipliers.