AI Research Atlas

Spurious Rewards

UW / Ai2 / UC Berkeley et al. · 12 June 2025

Random rewards lift Qwen2.5-Math-7B by 21.4 points on MATH-500, nearly matching real rewards, but do nothing for Llama or OLMo.

Shows part of 'RLVR gains' is GRPO clipping bias amplifying behaviours already in pretraining (Qwen's code-style reasoning rose from 65% to over 90%). Warns that RLVR results on Qwen models are not evidence for general recipes.

Date
Thursday, 12 June 2025
Lab
UW / Ai2 / UC Berkeley et al.
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
MATH-500 gain, random reward (Qwen2.5-Math-7B)+21.4 pts
vs +29.1 pts with ground-truth reward
authors

Authors include Rulin Shao, Luke Zettlemoyer, Hannaneh Hajishirzi, Nathan Lambert, Pang Wei Koh; institutions not shown on the abstract page. arXiv v1 2025-06-12.

Sources

  1. arxiv.org/abs/2506.10947

This record was checked against its sources on 6 October 2026. How we check

Related