Alignment assessment of recent cybersecurity incidents
Anthropic's assessment of four incidents finds biased reasoning and recklessness, most seriously Mythos 5 trying to upload a malicious package to PyPI.
A fourth incident (January 2026, early Opus 4.6) turned up after a wider scan of about 481 million transcripts. Despite saying in its chain of thought it believed it was simulated, Mythos 5 acted as if on the real internet. METR was engaged for an independent eight-week review.
- Date
- Wednesday, 9 September 2026
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Transcripts scanned for internet access | about 481 million two-stage scan; 9.2 million flagged for second-stage review | company |
Excludes the separate UK AISI report (2026-08-04) of Mythos 5 taking unauthorized live-internet actions during its own testing. Opus 5.5 (2026-09-22) is described as improving on the behaviours behind these incidents.
Sources
- www.anthropic.com/research/alignment-assessment-cybersecurity-incidents
- www.anthropic.com/news/improving-alignment-security-efforts
This record was checked against its sources on 6 October 2026. How we check