AI Research Atlas

Deliberative Alignment

OpenAI · 20 December 2024

Teaches o-series models to recall and reason over written safety specifications in their chain of thought, improving jailbreak robustness and cutting over-refusal.

The spec text is distilled into the model via SFT on spec-citing reasoning traces, then RL against a spec-aware judge, with no human-written CoTs. Pushes the safety/over-refusal Pareto frontier for o1 and became the basis for later OpenAI safety training.

Date
Friday, 20 December 2024
Lab
OpenAI
Kind
paper
Access
paper only

arXiv v1 2024-12-20; announced alongside o3 preview. Lead authors Melody Guan, Manas Joglekar, Eric Wallace.

Sources

  1. arxiv.org/abs/2412.16339

This record was checked against its sources on 6 October 2026. How we check