Spurious Rewards
Random rewards lift Qwen2.5-Math-7B by 21.4 points on MATH-500, nearly matching real rewards, but do nothing for Llama or OLMo.
Shows part of 'RLVR gains' is GRPO clipping bias amplifying behaviours already in pretraining (Qwen's code-style reasoning rose from 65% to over 90%). Warns that RLVR results on Qwen models are not evidence for general recipes.
- Date
- Thursday, 12 June 2025
- Lab
- UW / Ai2 / UC Berkeley et al.
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| MATH-500 gain, random reward (Qwen2.5-Math-7B) | +21.4 pts vs +29.1 pts with ground-truth reward | authors |
Authors include Rulin Shao, Luke Zettlemoyer, Hannaneh Hajishirzi, Nathan Lambert, Pang Wei Koh; institutions not shown on the abstract page. arXiv v1 2025-06-12.
Sources
This record was checked against its sources on 6 October 2026. How we check