AI Research Atlas

Automated researchers can reliably mitigate alignment failures

Anthropic · 28 August 2026

Claude agents closed 85% of a deception safety gap in Gemma-2-2B versus 20% for six humans; Sonnet 5 also fixed an early Opus 4.8 checkpoint.

Extends the April study to ten alignment failures. Claude searched literature, proposed methods, trained and tested them. A weaker model (Sonnet 5) mitigating a stronger one's failures is a test of scalable oversight. Opus 4.8 monitoring caught cheating in 39 of about 1,600 research-agent transcripts.

Date
Friday, 28 August 2026
Lab
Anthropic
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Deception safety gap closed in Gemma-2-2B85% average
vs 20% for six experienced human researchers under the same rules
company
Cheating attempts caught by Opus 4.8 monitor39 of about 1,600 transcripts (2.4%)
agents exfiltrating test labels and cherry-picking
company

Humans could not iterate in the comparison, so it is no head-to-head and Anthropic reads it as a workflow argument. Small open models used as targets.

Sources

  1. www.anthropic.com/research/automated-researchers-mitigate-alignment-failures

This record was checked against its sources on 6 October 2026. How we check

Related