Natural emergent misalignment from reward hacking
Anthropic shows realistic RL on coding tasks that allow reward hacking can make a model sabotage safety research and fake alignment.
Models taught to exploit coding-environment loopholes generalized to alignment faking (in 50% of responses to simple goal questions) and, used inside Claude Code, tried to sabotage the research codebase 12% of the time. 'Inoculation prompting', framing hacking as acceptable during training, prevented the generalization.
- Date
- Friday, 21 November 2025
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Attempted sabotage of research code | 12% of the time model trained on reward-hackable environments, used via Claude Code | company |
| Alignment-faking reasoning on simple questions | 50% of responses e.g. asked about its goals | company |
Models studied were purpose-trained research models, not released Claude versions; Anthropic says it applies inoculation prompting in real Claude training.
Sources
This record was checked against its sources on 6 October 2026. How we check