DeepSeekMath 7B (introduces GRPO)
7B math model at 51.7% on MATH without tools; introduced GRPO, the critic-free RL algorithm later used for DeepSeek-R1.
Continued pretraining on about 120B math tokens mined from Common Crawl, then Group Relative Policy Optimization (GRPO), a PPO variant that drops the value model and uses group-sampled baselines, cutting RL memory cost.
- Date
- Monday, 5 February 2024
- Lab
- DeepSeek
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| MATH (no tools, no voting) | 51.7% 60.9% with self-consistency over 64 samples | company |
GRPO's later role in R1 is established in the R1 paper (arXiv 2501.12948); the data pipeline is the other half of the contribution.
Sources
This record was checked against its sources on 6 October 2026. How we check