Reasoning models struggle to control their chains of thought
OpenAI finds frontier reasoning models rarely control their chain of thought even when told it is monitored, which supports CoT monitoring.
Introduces CoT controllability as a measure of whether a model can shape its reasoning trace to evade monitors; controllability is low across frontier models and higher for simple tasks. Later system cards track it, and GPT-6 Astra shows a decline in monitorability.
- Date
- Thursday, 5 March 2026
- Lab
- OpenAI
- Kind
- paper
- Access
- research preview
Builds on OpenAI's 2025-03 CoT monitoring work and the 2025-07 multi-lab monitorability position paper (both in the papers dataset).
Sources
- openai.com/index/reasoning-models-chain-of-thought-controllability/
- deploymentsafety.openai.com/gpt-6-astra/
This record was checked against its sources on 6 October 2026. How we check