AI Research Atlas

s1: Simple test-time scaling

Stanford / UW / Ai2 · 31 January 2025

1,000 curated reasoning traces plus 'budget forcing' (appending 'Wait') turn Qwen2.5-32B into an o1-preview-beating math model.

Shows a tiny SFT set (s1K) elicits long reasoning, and that forcibly extending or truncating the thinking trace gives a clean test-time scaling curve (AIME24 50% to 57%). Everything was released open, making it the cheapest reproducible reasoning recipe.

Date
Friday, 31 January 2025
Lab
Stanford / UW / Ai2
Kind
paper
Access
open weights

Figures

MeasureValueMeasured by
Margin vs o1-preview on competition mathup to +27%
MATH and AIME24, s1-32B
authors
AIME24 with budget forcing50% -> 57%
extrapolation beyond the trained length
authors
SFT examples1,000
s1K
authors

Lead author Niklas Muennighoff. arXiv v1 2025-01-31. The reasoning traces for s1K were distilled from a proprietary thinking model (Gemini), so it is distillation-assisted; this attribution comes from the paper and was not re-opened on this pass.

Sources

  1. arxiv.org/abs/2501.19393

This record was checked against its sources on 6 October 2026. How we check

Related