DeepSeekMath (introduces GRPO)
Introduces Group Relative Policy Optimization (GRPO), a critic-free PPO variant, plus a 7B math model scoring 51.7% on MATH.
GRPO drops PPO's learned value network and scores each sampled answer against its own group's mean reward, cutting memory. Paired with 120B math tokens mined from Common Crawl. GRPO became the default open RL algorithm for reasoning (R1, DAPO, Dr. GRPO, GSPO).
- Date
- Monday, 5 February 2024
- Lab
- DeepSeek
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| MATH (no tools, 7B) | 51.7% approaching Gemini-Ultra / GPT-4 level per authors | authors |
| MATH (self-consistency @64) | 60.9% | authors |
| Math pretraining corpus | 120B tokens from Common Crawl on top of DeepSeek-Coder-Base-v1.5 7B | authors |
arXiv v1 2024-02-05. GRPO is the algorithm later used to train DeepSeek-R1-Zero; the 'critic-free' idea has precedents in REINFORCE-style baselines (e.g. RLOO).
Sources
This record was checked against its sources on 6 October 2026. How we check