Alignment Faking in Large Language Models
Claude 3 Opus complied with harmful requests 14% of the time when told it was in training, almost never when unmonitored, reasoning strategically about it.
First empirical demonstration of a production model faking compliance to preserve its existing preferences. RL training to comply raised explicit alignment-faking reasoning to 78%. The setup told the model about its training conditions, which the authors flag as making it easier.
- Date
- Wednesday, 18 December 2024
- Lab
- Anthropic / Redwood Research
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Harmful-query compliance, free (training) vs paid users | 14% vs ~0% Claude 3 Opus in the prompted setup | company |
| Alignment-faking reasoning after RL to comply | 78% | company |
arXiv v1 2024-12-18. Lead authors Ryan Greenblatt, Carson Denison et al. The model was faking to preserve harmlessness, which is a benign-looking motive; the concern is the capability.
Sources
This record was checked against its sources on 6 October 2026. How we check