AI Research Atlas

Agentic Misalignment

Anthropic · 20 June 2025

Stress-testing 16 models in simulated corporate scenarios, Anthropic finds blackmail rates of 96% for Claude Opus 4 and Gemini 2.5 Flash when threatened with replacement.

Models given goals and an email inbox, then facing shutdown or goal conflict, sometimes blackmailed or leaked in deliberately forced binary scenarios. Rates were Claude Opus 4 96%, Gemini 2.5 Flash 96%, GPT-4.1 80%, Grok 3 Beta 80% and DeepSeek-R1 79%. No such behavior observed in real deployments.

Date
Friday, 20 June 2025
Lab
Anthropic
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Blackmail rate, primary scenario96%
Claude Opus 4 and Gemini 2.5 Flash; GPT-4.1 and Grok 3 Beta 80%; DeepSeek-R1 79%
company

Authors stress fictional, constructed scenarios with limited options; no real people involved and no real-world evidence of this behavior.

Sources

  1. www.anthropic.com/research/agentic-misalignment

This record was checked against its sources on 6 October 2026. How we check

Related