AI Research Atlas

Large Language Monkeys (repeated sampling)

Stanford / Google DeepMind / Oxford · 31 July 2024

Coverage (any-sample-correct) scales log-linearly with samples over four orders of magnitude; SWE-bench Lite goes from 15.9% to 56% with 250 samples.

Treats the number of samples per problem as a scaling axis. With automatic verifiers (unit tests, proof checkers) more samples keep solving more problems; without verifiers, majority vote and reward models plateau, making verification the bottleneck.

Date
Wednesday, 31 July 2024
Lab
Stanford / Google DeepMind / Oxford
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
SWE-bench Lite (DeepSeek-Coder-V2-Instruct)15.9% -> 56%
1 sample vs 250 samples; prior single-sample SOTA 43%
authors

Affiliation line is from the author list (Ré, Mirhoseini at Stanford; Le at Google; Clark at Oxford); the abstract page itself lists names only.

Sources

  1. arxiv.org/abs/2407.21787

This record was checked against its sources on 6 October 2026. How we check