ProRL: Prolonged RL
Argues that prolonged RL with KL control and reference resets expands reasoning boundaries beyond the base model, answering the pass@k critique.
Runs RL for far longer on diverse tasks, with KL regularisation and periodic reference-policy resets. RL models beat the base even at large k on tasks where the base fails entirely, so gains depend on training duration and base competence. Released a 1.5B reasoning model.
- Date
- Friday, 30 May 2025
- Lab
- NVIDIA
- Kind
- paper
- Access
- open weights
Authors incl. Mingjie Liu, Yejin Choi, Jan Kautz, Yi Dong. The weights are Nemotron-Research-Reasoning-Qwen-1.5B.
Sources
This record was checked against its sources on 6 October 2026. How we check