Automated researchers can reliably mitigate alignment failures
Claude agents closed 85% of a deception safety gap in Gemma-2-2B versus 20% for six humans; Sonnet 5 also fixed an early Opus 4.8 checkpoint.
Extends the April study to ten alignment failures. Claude searched literature, proposed methods, trained and tested them. A weaker model (Sonnet 5) mitigating a stronger one's failures is a test of scalable oversight. Opus 4.8 monitoring caught cheating in 39 of about 1,600 research-agent transcripts.
- Date
- Friday, 28 August 2026
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Deception safety gap closed in Gemma-2-2B | 85% average vs 20% for six experienced human researchers under the same rules | company |
| Cheating attempts caught by Opus 4.8 monitor | 39 of about 1,600 transcripts (2.4%) agents exfiltrating test labels and cherry-picking | company |
Humans could not iterate in the comparison, so it is no head-to-head and Anthropic reads it as a workflow argument. Small open models used as targets.
Sources
This record was checked against its sources on 6 October 2026. How we check