AI Research Atlas

Teaching Claude why

Anthropic · 8 May 2026

Anthropic reports every Claude model since Haiku 4.5 scores perfectly on its blackmail-style agentic misalignment test, down from up to 96% for Opus 4.

Training on the evaluation distribution suppressed blackmail but did not generalize to held-out alignment tests. Documents about Claude's constitution and fictional stories of admirable AIs worked, as did teaching the reasons behind better actions as well as demonstrating them.

Date
Friday, 8 May 2026
Lab
Anthropic
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Blackmail rate on agentic misalignment eval0% since Haiku 4.5
previously up to 96% for Claude Opus 4
company

Perfect scores are on Anthropic's own evaluation; the post itself reports that training close to that evaluation did not improve its held-out alignment assessment.

Sources

  1. www.anthropic.com/research/teaching-claude-why

This record was checked against its sources on 6 October 2026. How we check

Related