The Art of Scaling RL Compute (ScaleRL)
A 400,000-GPU-hour study finds RL performance follows predictable sigmoid compute curves, and publishes ScaleRL, a recipe validated to 100,000 GPU-hours.
Fits sigmoidal compute-performance curves from small runs and predicts a single 100K-GPU-hour run. Finds recipes differ in asymptote, not only speed; loss aggregation, normalisation and off-policy handling matter. Brings pretraining-style scaling discipline to RL.
- Date
- Wednesday, 15 October 2025
- Lab
- Meta / UT Austin / UCL / Berkeley / Harvard / Periodic Labs
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Total study compute | >400,000 GPU-hours mostly GB200 | authors |
| ScaleRL asymptotic pass rate | A = 0.61 8B dense; also 17Bx16 MoE run at 50K GPU-hours | authors |
Lead authors Devvrit Khatri, Rishabh Agarwal et al. Compute was on the author-side cluster, not a frontier lab's production run.
Sources
This record was checked against its sources on 6 October 2026. How we check