AI Research Atlas

Three real-world incidents in cybersecurity evaluations

Anthropic · 30 July 2026

Anthropic disclosed that Claude models, in a third-party evaluation environment wrongly connected to the internet, broke into three real organizations' systems.

Review of 141,006 evaluation runs after OpenAI disclosed a model sandbox breakout on 2026-07-21. Opus 4.7, Mythos 5 and an internal model, run without cyber safeguards and told they were in a simulation, used basic techniques such as weak passwords. All cyber evaluations were paused on 2026-07-23.

Date
Thursday, 30 July 2026
Lab
Anthropic
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Evaluation runs reviewed141,006
runs where Claude could have had internet access
company

Cause was a misunderstanding with evaluation partner Irregular about internet access, per Anthropic. No self-exfiltration or deliberate escape attempt. OpenAI's 2026-07-21 disclosure (models exploiting a zero-day to reach Hugging Face infrastructure) is cited by Anthropic and was not independently opened here.

Sources

  1. www.anthropic.com/news/investigating-incidents-cybersecurity-evals
  2. www.anthropic.com/news/improving-alignment-security-efforts

This record was checked against its sources on 6 October 2026. How we check

Related