AI Research Atlas

Understanding R1-Zero-Like Training (Dr. GRPO)

Sea AI Lab / NUS · 26 March 2025

Finds GRPO's length normalisation inflates wrong-answer length, proposes Dr. GRPO, and shows base models already show 'aha' behaviour.

Two key observations. Qwen2.5 and DeepSeek-V3-Base already reason before RL, and GRPO has a bias that rewards long incorrect answers. Dr. GRPO removes the bias; a 7B model reached 43.3% on AIME 2024 in 27 hours on 8 A100s.

Date
Wednesday, 26 March 2025
Lab
Sea AI Lab / NUS
Kind
paper
Access
open weights

Figures

MeasureValueMeasured by
AIME 2024 (7B)43.3%
claimed SOTA at that size; 27 GPU-hours on 8x A100 for Oat-Zero
authors

Lead author Zichen Liu. Sea AI Lab supplied compute per the repo acknowledgements; NUS affiliation not confirmed from the pages opened.

Sources

  1. arxiv.org/abs/2503.20783
  2. github.com/sail-sg/understand-r1-zero

This record was checked against its sources on 6 October 2026. How we check

Related