Deliberative Alignment
Teaches o-series models to recall and reason over written safety specifications in their chain of thought, improving jailbreak robustness and cutting over-refusal.
The spec text is distilled into the model via SFT on spec-citing reasoning traces, then RL against a spec-aware judge, with no human-written CoTs. Pushes the safety/over-refusal Pareto frontier for o1 and became the basis for later OpenAI safety training.
- Date
- Friday, 20 December 2024
- Lab
- OpenAI
- Kind
- paper
- Access
- paper only
arXiv v1 2024-12-20; announced alongside o3 preview. Lead authors Melody Guan, Manas Joglekar, Eric Wallace.
Sources
This record was checked against its sources on 6 October 2026. How we check