Teaching Claude why
Anthropic reports every Claude model since Haiku 4.5 scores perfectly on its blackmail-style agentic misalignment test, down from up to 96% for Opus 4.
Training on the evaluation distribution suppressed blackmail but did not generalize to held-out alignment tests. Documents about Claude's constitution and fictional stories of admirable AIs worked, as did teaching the reasons behind better actions as well as demonstrating them.
- Date
- Friday, 8 May 2026
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Blackmail rate on agentic misalignment eval | 0% since Haiku 4.5 previously up to 96% for Claude Opus 4 | company |
Perfect scores are on Anthropic's own evaluation; the post itself reports that training close to that evaluation did not improve its held-out alignment assessment.
Sources
This record was checked against its sources on 6 October 2026. How we check