AI Research Atlas

Monitoring Reasoning Models for Misbehavior

OpenAI · 14 March 2025

A weaker GPT-4o can catch o3-mini reward hacking from its CoT, but training against the monitor teaches the model to hide intent.

Shows CoT monitoring beats action-only monitoring at detecting reward hacks in coding RL, then shows 'obfuscated reward hacking' when CoT is optimised against. Introduces the 'monitorability tax', which means leaving the CoT unpressured.

Date
Friday, 14 March 2025
Lab
OpenAI
Kind
paper
Access
paper only

arXiv v1 2025-03-14. Lead author Bowen Baker. Became the policy reference for later multi-lab 'do not train on the CoT' commitments.

Sources

  1. arxiv.org/abs/2503.11926

This record was checked against its sources on 6 October 2026. How we check