Reinforcement Pre-Training (RPT)
Recasts next-token prediction as a reasoning task with verifiable reward (did the token match?), scaling RL over ordinary web text.
A 14B model trained with RPT reaches 45.1% next-token accuracy on easy splits versus 41.6% for the same-size distilled baseline, and gives a stronger starting point for later RL. Points to RL on unlabeled text rather than only curated verifiable tasks.
- Date
- Monday, 9 June 2025
- Lab
- Microsoft Research / Peking / Tsinghua
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Next-token accuracy, easy split (14B) | 45.11% vs 41.60% RPT-14B vs R1-Distill-Qwen-14B | authors |
Preliminary, small-scale; no frontier lab has confirmed adopting it.
Sources
This record was checked against its sources on 6 October 2026. How we check