Large Language Monkeys (repeated sampling)
Coverage (any-sample-correct) scales log-linearly with samples over four orders of magnitude; SWE-bench Lite goes from 15.9% to 56% with 250 samples.
Treats the number of samples per problem as a scaling axis. With automatic verifiers (unit tests, proof checkers) more samples keep solving more problems; without verifiers, majority vote and reward models plateau, making verification the bottleneck.
- Date
- Wednesday, 31 July 2024
- Lab
- Stanford / Google DeepMind / Oxford
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| SWE-bench Lite (DeepSeek-Coder-V2-Instruct) | 15.9% -> 56% 1 sample vs 250 samples; prior single-sample SOTA 43% | authors |
Affiliation line is from the author list (Ré, Mirhoseini at Stanford; Le at Google; Clark at Oxford); the abstract page itself lists names only.
Sources
This record was checked against its sources on 6 October 2026. How we check