Scaling LLM Test-Time Compute Optimally (Snell et al.)
Shows that allocating inference compute per prompt difficulty can beat a 14x larger model, giving the first rigorous test-time-scaling recipe.
Compares verifier-guided search against adaptive revision of the model's own answers, and picks the strategy by question difficulty. A compute-optimal policy gives over 4x the efficiency of best-of-N, and on easier prompts a small model with extra inference compute matches a model 14x larger in pretraining FLOPs.
- Date
- Tuesday, 6 August 2024
- Lab
- UC Berkeley / Google DeepMind
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Efficiency vs best-of-N baseline | >4x compute-optimal allocation | authors |
| Small model + test-time compute vs larger model | matches 14x larger on problems where small model has non-trivial success rate | authors |
Appeared about five weeks before OpenAI's o1-preview (2024-09-12), which popularised the same idea in a product. Authors are Charlie Snell, Jaehoon Lee, Kelvin Xu and Aviral Kumar.
Sources
This record was checked against its sources on 6 October 2026. How we check