Understanding R1-Zero-Like Training (Dr. GRPO)
Finds GRPO's length normalisation inflates wrong-answer length, proposes Dr. GRPO, and shows base models already show 'aha' behaviour.
Two key observations. Qwen2.5 and DeepSeek-V3-Base already reason before RL, and GRPO has a bias that rewards long incorrect answers. Dr. GRPO removes the bias; a 7B model reached 43.3% on AIME 2024 in 27 hours on 8 A100s.
- Date
- Wednesday, 26 March 2025
- Lab
- Sea AI Lab / NUS
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| AIME 2024 (7B) | 43.3% claimed SOTA at that size; 27 GPU-hours on 8x A100 for Oat-Zero | authors |
Lead author Zichen Liu. Sea AI Lab supplied compute per the repo acknowledgements; NUS affiliation not confirmed from the pages opened.
Sources
This record was checked against its sources on 6 October 2026. How we check