AI Research Atlas

DeepSeekMath 7B (introduces GRPO)

DeepSeek · 5 February 2024

7B math model at 51.7% on MATH without tools; introduced GRPO, the critic-free RL algorithm later used for DeepSeek-R1.

Continued pretraining on about 120B math tokens mined from Common Crawl, then Group Relative Policy Optimization (GRPO), a PPO variant that drops the value model and uses group-sampled baselines, cutting RL memory cost.

Date
Monday, 5 February 2024
Lab
DeepSeek
Kind
open-weights
Access
open weights (restricted license)

Figures

MeasureValueMeasured by
MATH (no tools, no voting)51.7%
60.9% with self-consistency over 64 samples
company

GRPO's later role in R1 is established in the R1 paper (arXiv 2501.12948); the data pipeline is the other half of the contribution.

Sources

  1. arxiv.org/abs/2402.03300
  2. huggingface.co/api/models/deepseek-ai/deepseek-math-7b-base

This record was checked against its sources on 6 October 2026. How we check