DAPO
Fully open large-scale RL system reaches 50 on AIME 2024 with Qwen2.5-32B, beating R1-Zero-Qwen-32B in half the training steps.
Four fixes to GRPO: Clip-Higher (stops entropy collapse), Dynamic Sampling (drops all-correct/all-wrong prompts), token-level loss and overlong reward shaping. Released code, data and recipe to close the gap left by closed details in R1 and o1.
- Date
- Tuesday, 18 March 2025
- Lab
- ByteDance Seed / Tsinghua AIR
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| AIME 2024 | 50 Qwen2.5-32B base; ~50% fewer steps than R1-Zero-Qwen-32B | authors |
Lead author Qiying Yu; affiliations ByteDance Seed, Tsinghua AIR, HKU, SIA-Lab. arXiv v1 2025-03-18, v2 2025-05-20.
Sources
This record was checked against its sources on 6 October 2026. How we check