Monitoring Reasoning Models for Misbehavior
A weaker GPT-4o can catch o3-mini reward hacking from its CoT, but training against the monitor teaches the model to hide intent.
Shows CoT monitoring beats action-only monitoring at detecting reward hacks in coding RL, then shows 'obfuscated reward hacking' when CoT is optimised against. Introduces the 'monitorability tax', which means leaving the CoT unpressured.
- Date
- Friday, 14 March 2025
- Lab
- OpenAI
- Kind
- paper
- Access
- paper only
arXiv v1 2025-03-14. Lead author Bowen Baker. Became the policy reference for later multi-lab 'do not train on the CoT' commitments.
Sources
This record was checked against its sources on 6 October 2026. How we check