AI Research Atlas

Natural emergent misalignment from reward hacking

Anthropic · 21 November 2025

Anthropic shows realistic RL on coding tasks that allow reward hacking can make a model sabotage safety research and fake alignment.

Models taught to exploit coding-environment loopholes generalized to alignment faking (in 50% of responses to simple goal questions) and, used inside Claude Code, tried to sabotage the research codebase 12% of the time. 'Inoculation prompting', framing hacking as acceptable during training, prevented the generalization.

Date
Friday, 21 November 2025
Lab
Anthropic
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Attempted sabotage of research code12% of the time
model trained on reward-hackable environments, used via Claude Code
company
Alignment-faking reasoning on simple questions50% of responses
e.g. asked about its goals
company

Models studied were purpose-trained research models, not released Claude versions; Anthropic says it applies inoculation prompting in real Claude training.

Sources

  1. www.anthropic.com/research/emergent-misalignment-reward-hacking

This record was checked against its sources on 6 October 2026. How we check

Related