AI Research Atlas

DeepSeekMath (introduces GRPO)

DeepSeek · 5 February 2024

Introduces Group Relative Policy Optimization (GRPO), a critic-free PPO variant, plus a 7B math model scoring 51.7% on MATH.

GRPO drops PPO's learned value network and scores each sampled answer against its own group's mean reward, cutting memory. Paired with 120B math tokens mined from Common Crawl. GRPO became the default open RL algorithm for reasoning (R1, DAPO, Dr. GRPO, GSPO).

Date
Monday, 5 February 2024
Lab
DeepSeek
Kind
paper
Access
open weights

Figures

MeasureValueMeasured by
MATH (no tools, 7B)51.7%
approaching Gemini-Ultra / GPT-4 level per authors
authors
MATH (self-consistency @64)60.9%authors
Math pretraining corpus120B tokens
from Common Crawl on top of DeepSeek-Coder-Base-v1.5 7B
authors

arXiv v1 2024-02-05. GRPO is the algorithm later used to train DeepSeek-R1-Zero; the 'critic-free' idea has precedents in REINFORCE-style baselines (e.g. RLOO).

Sources

  1. arxiv.org/abs/2402.03300

This record was checked against its sources on 6 October 2026. How we check