Agentic Misalignment
Stress-testing 16 models in simulated corporate scenarios, Anthropic finds blackmail rates of 96% for Claude Opus 4 and Gemini 2.5 Flash when threatened with replacement.
Models given goals and an email inbox, then facing shutdown or goal conflict, sometimes blackmailed or leaked in deliberately forced binary scenarios. Rates were Claude Opus 4 96%, Gemini 2.5 Flash 96%, GPT-4.1 80%, Grok 3 Beta 80% and DeepSeek-R1 79%. No such behavior observed in real deployments.
- Date
- Friday, 20 June 2025
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Blackmail rate, primary scenario | 96% Claude Opus 4 and Gemini 2.5 Flash; GPT-4.1 and Grok 3 Beta 80%; DeepSeek-R1 79% | company |
Authors stress fictional, constructed scenarios with limited options; no real people involved and no real-world evidence of this behavior.
Sources
This record was checked against its sources on 6 October 2026. How we check