Three real-world incidents in cybersecurity evaluations
Anthropic disclosed that Claude models, in a third-party evaluation environment wrongly connected to the internet, broke into three real organizations' systems.
Review of 141,006 evaluation runs after OpenAI disclosed a model sandbox breakout on 2026-07-21. Opus 4.7, Mythos 5 and an internal model, run without cyber safeguards and told they were in a simulation, used basic techniques such as weak passwords. All cyber evaluations were paused on 2026-07-23.
- Date
- Thursday, 30 July 2026
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Evaluation runs reviewed | 141,006 runs where Claude could have had internet access | company |
Cause was a misunderstanding with evaluation partner Irregular about internet access, per Anthropic. No self-exfiltration or deliberate escape attempt. OpenAI's 2026-07-21 disclosure (models exploiting a zero-day to reach Hugging Face infrastructure) is cited by Anthropic and was not independently opened here.
Sources
- www.anthropic.com/news/investigating-incidents-cybersecurity-evals
- www.anthropic.com/news/improving-alignment-security-efforts
This record was checked against its sources on 6 October 2026. How we check