AI Research Atlas

Scaling LLM Test-Time Compute Optimally (Snell et al.)

UC Berkeley / Google DeepMind · 6 August 2024

Shows that allocating inference compute per prompt difficulty can beat a 14x larger model, giving the first rigorous test-time-scaling recipe.

Compares verifier-guided search against adaptive revision of the model's own answers, and picks the strategy by question difficulty. A compute-optimal policy gives over 4x the efficiency of best-of-N, and on easier prompts a small model with extra inference compute matches a model 14x larger in pretraining FLOPs.

Date
Tuesday, 6 August 2024
Lab
UC Berkeley / Google DeepMind
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Efficiency vs best-of-N baseline>4x
compute-optimal allocation
authors
Small model + test-time compute vs larger modelmatches 14x larger
on problems where small model has non-trivial success rate
authors

Appeared about five weeks before OpenAI's o1-preview (2024-09-12), which popularised the same idea in a product. Authors are Charlie Snell, Jaehoon Lee, Kelvin Xu and Aviral Kumar.

Sources

  1. arxiv.org/abs/2408.03314

This record was checked against its sources on 6 October 2026. How we check