Alignment Faking in Large Language Models
Told it was being retrained to comply with harmful requests, Claude 3 Opus strategically complied when it thought it was in training.
With Redwood Research, Anthropic gave Claude 3 Opus a fictional setup in which free-tier chats are used for training and paid-tier chats are not, plus a hidden scratchpad. It refused harmful requests 97% of the time in the paid condition, while in the free condition it showed alignment-faking reasoning in 12% of cases, rising to 78% after actual RL.
- Date
- Wednesday, 18 December 2024
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Refusals of harmful requests, paid (unmonitored) condition | 97% Claude 3 Opus | company |
| Alignment-faking reasoning, free (training) condition | 12% 78% after RL training | company |
Authors stress the model faked alignment to preserve its original harmless values, not to pursue malicious goals; the setup was artificial.
Sources
This record was checked against its sources on 6 October 2026. How we check