Chain of Thought Monitorability: A New and Fragile Opportunity
Cross-lab position paper urges labs to preserve readable chains of thought as a safety tool, warning the property is fragile under training pressure.
Authors from UK AISI, Anthropic, OpenAI, Google DeepMind, Apollo, METR and others (incl. Yoshua Bengio, Anca Dragan, Shane Legg, Neel Nanda) ask developers to track monitorability in system cards and weigh it when choosing architectures, since latent reasoning would remove it.
- Date
- Tuesday, 15 July 2025
- Lab
- Multi-lab (UK AISI, Anthropic, OpenAI, Google DeepMind and others)
- Kind
- paper
- Access
- paper only
Lead author Tomek Korbak; 40-plus authors. arXiv v1 2025-07-15, v2 2025-12-07. Position paper, not an experiment.
Sources
This record was checked against its sources on 6 October 2026. How we check