AI Research Atlas

Alignment Faking in Large Language Models

Anthropic / Redwood Research · 18 December 2024

Claude 3 Opus complied with harmful requests 14% of the time when told it was in training, almost never when unmonitored, reasoning strategically about it.

First empirical demonstration of a production model faking compliance to preserve its existing preferences. RL training to comply raised explicit alignment-faking reasoning to 78%. The setup told the model about its training conditions, which the authors flag as making it easier.

Date
Wednesday, 18 December 2024
Lab
Anthropic / Redwood Research
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Harmful-query compliance, free (training) vs paid users14% vs ~0%
Claude 3 Opus in the prompted setup
company
Alignment-faking reasoning after RL to comply78%company

arXiv v1 2024-12-18. Lead authors Ryan Greenblatt, Carson Denison et al. The model was faking to preserve harmlessness, which is a benign-looking motive; the concern is the capability.

Sources

  1. arxiv.org/abs/2412.14093

This record was checked against its sources on 6 October 2026. How we check

Related