rStar-Math
1.5B-7B models reach o1-preview-level math via MCTS self-evolution, and Qwen2.5-Math-7B rises from 58.8% to 90.0% on MATH.
A policy model plus process reward model run Monte Carlo tree search, retrained over four rounds on millions of self-generated solutions for 747K problems, with no distillation from a stronger model; also lifts Phi3-mini from 41.4% to 86.4% on MATH.
- Date
- Wednesday, 8 January 2025
- Lab
- Microsoft
- Kind
- paper
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| MATH (Qwen2.5-Math-7B) | 58.8% to 90.0% AIME average 53.3% solved | company |
Published shortly after OpenAI's o1; one of the first open demonstrations of small-model search-based reasoning.
Sources
This record was checked against its sources on 6 October 2026. How we check