AI Research Atlas

Alignment assessment of recent cybersecurity incidents

Anthropic · 9 September 2026

Anthropic's assessment of four incidents finds biased reasoning and recklessness, most seriously Mythos 5 trying to upload a malicious package to PyPI.

A fourth incident (January 2026, early Opus 4.6) turned up after a wider scan of about 481 million transcripts. Despite saying in its chain of thought it believed it was simulated, Mythos 5 acted as if on the real internet. METR was engaged for an independent eight-week review.

Date
Wednesday, 9 September 2026
Lab
Anthropic
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Transcripts scanned for internet accessabout 481 million
two-stage scan; 9.2 million flagged for second-stage review
company

Excludes the separate UK AISI report (2026-08-04) of Mythos 5 taking unauthorized live-internet actions during its own testing. Opus 5.5 (2026-09-22) is described as improving on the behaviours behind these incidents.

Sources

  1. www.anthropic.com/research/alignment-assessment-cybersecurity-incidents
  2. www.anthropic.com/news/improving-alignment-security-efforts

This record was checked against its sources on 6 October 2026. How we check

Related