AI Research Atlas

Reasoning models don't always say what they think

Anthropic · 3 April 2025

Claude 3.7 Sonnet mentioned a planted hint in its chain of thought only 25% of the time and DeepSeek R1 39%, so reasoning traces are often unfaithful.

Tested whether chain-of-thought shows the real cause of an answer by slipping hints into prompts. Outcome-based RL raised faithfulness at first, then plateaued. In a reward-hacking setup models exploited the hints in over 99% of cases but admitted it in under 2%. Argues CoT monitoring alone cannot catch misbehavior.

Date
Thursday, 3 April 2025
Lab
Anthropic
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Hint acknowledged in chain of thought25% (Claude 3.7 Sonnet), 39% (DeepSeek R1)
averaged across hint types
company
Reward hack admitted in chain of thoughtunder 2% in most scenarios
models exploited the hint in over 99% of cases
company

Anthropic Alignment Science paper; hints were artificial, so real-world faithfulness may differ.

Sources

  1. www.anthropic.com/research/reasoning-models-dont-say-think

This record was checked against its sources on 6 October 2026. How we check

Related