s1: Simple test-time scaling
1,000 curated reasoning traces plus 'budget forcing' (appending 'Wait') turn Qwen2.5-32B into an o1-preview-beating math model.
Shows a tiny SFT set (s1K) elicits long reasoning, and that forcibly extending or truncating the thinking trace gives a clean test-time scaling curve (AIME24 50% to 57%). Everything was released open, making it the cheapest reproducible reasoning recipe.
- Date
- Friday, 31 January 2025
- Lab
- Stanford / UW / Ai2
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Margin vs o1-preview on competition math | up to +27% MATH and AIME24, s1-32B | authors |
| AIME24 with budget forcing | 50% -> 57% extrapolation beyond the trained length | authors |
| SFT examples | 1,000 s1K | authors |
Lead author Niklas Muennighoff. arXiv v1 2025-01-31. The reasoning traces for s1K were distilled from a proprietary thinking model (Gemini), so it is distillation-assisted; this attribution comes from the paper and was not re-opened on this pass.
Sources
This record was checked against its sources on 6 October 2026. How we check