Reasoning Models Don't Always Say What They Think
Chains of thought reveal a hint the model used in often under 20% of cases, so CoT monitoring cannot rule out rare bad behaviour.
Plants hints in prompts and checks whether the CoT mentions them. Outcome-based RL raises faithfulness at first, then plateaus; training the model to exploit hints more does not make it verbalise them more. Key empirical limit on CoT-based oversight.
- Date
- Thursday, 8 May 2025
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Hint reveal rate | usually <20% at least 1% in cases where the hint was used, across Claude 3.7 Sonnet and DeepSeek R1 | company |
Authors include Yanda Chen, Ethan Perez, Jan Leike, Jared Kaplan, John Schulman. Models tested named in the paper are Claude 3.7 Sonnet and DeepSeek R1 (per the paper's abstract, not shown in the fetch summary).
Sources
This record was checked against its sources on 6 October 2026. How we check