Does RL Really Incentivize Reasoning Beyond the Base Model?
RLVR models win at pass@1 but base models win at large pass@k, suggesting RL sharpens existing ability rather than creating new reasoning paths.
Measures pass@k up to large k across math, code and vision. Six RLVR algorithms perform similarly and far below optimal; distillation, unlike RL, did add new patterns. Set the terms of the 'does RL expand capability' debate. NeurIPS 2025 Best Paper Runner-Up.
- Date
- Friday, 18 April 2025
- Lab
- Tsinghua LeapLab / SJTU
- Kind
- paper
- Access
- paper only
Lead author Yang Yue. The finding is contested. ProRL (NVIDIA, 2025-05-30) argues longer RL does expand boundaries, and the answer depends on training length, base model and task.
Sources
This record was checked against its sources on 6 October 2026. How we check