AI Research Atlas

rStar-Math

Microsoft · 8 January 2025

1.5B-7B models reach o1-preview-level math via MCTS self-evolution, and Qwen2.5-Math-7B rises from 58.8% to 90.0% on MATH.

A policy model plus process reward model run Monte Carlo tree search, retrained over four rounds on millions of self-generated solutions for 747K problems, with no distillation from a stronger model; also lifts Phi3-mini from 41.4% to 86.4% on MATH.

Date
Wednesday, 8 January 2025
Lab
Microsoft
Kind
paper
Access
research preview

Figures

MeasureValueMeasured by
MATH (Qwen2.5-Math-7B)58.8% to 90.0%
AIME average 53.3% solved
company

Published shortly after OpenAI's o1; one of the first open demonstrations of small-model search-based reasoning.

Sources

  1. arxiv.org/abs/2501.04519

This record was checked against its sources on 6 October 2026. How we check