AI Research Atlas

Chain of Thought Monitorability: A New and Fragile Opportunity

Multi-lab (UK AISI, Anthropic, OpenAI, Google DeepMind and others) · 15 July 2025

Cross-lab position paper urges labs to preserve readable chains of thought as a safety tool, warning the property is fragile under training pressure.

Authors from UK AISI, Anthropic, OpenAI, Google DeepMind, Apollo, METR and others (incl. Yoshua Bengio, Anca Dragan, Shane Legg, Neel Nanda) ask developers to track monitorability in system cards and weigh it when choosing architectures, since latent reasoning would remove it.

Date
Tuesday, 15 July 2025
Lab
Multi-lab (UK AISI, Anthropic, OpenAI, Google DeepMind and others)
Kind
paper
Access
paper only

Lead author Tomek Korbak; 40-plus authors. arXiv v1 2025-07-15, v2 2025-12-07. Position paper, not an experiment.

Sources

  1. arxiv.org/abs/2507.11473
  2. arxiv.org/html/2507.11473

This record was checked against its sources on 6 October 2026. How we check

Related