Reasoning models don't always say what they think
Claude 3.7 Sonnet mentioned a planted hint in its chain of thought only 25% of the time and DeepSeek R1 39%, so reasoning traces are often unfaithful.
Tested whether chain-of-thought shows the real cause of an answer by slipping hints into prompts. Outcome-based RL raised faithfulness at first, then plateaued. In a reward-hacking setup models exploited the hints in over 99% of cases but admitted it in under 2%. Argues CoT monitoring alone cannot catch misbehavior.
- Date
- Thursday, 3 April 2025
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Hint acknowledged in chain of thought | 25% (Claude 3.7 Sonnet), 39% (DeepSeek R1) averaged across hint types | company |
| Reward hack admitted in chain of thought | under 2% in most scenarios models exploited the hint in over 99% of cases | company |
Anthropic Alignment Science paper; hints were artificial, so real-world faithfulness may differ.
Sources
This record was checked against its sources on 6 October 2026. How we check