AI Research Atlas

Reinforcement Pre-Training (RPT)

Microsoft Research / Peking / Tsinghua · 9 June 2025

Recasts next-token prediction as a reasoning task with verifiable reward (did the token match?), scaling RL over ordinary web text.

A 14B model trained with RPT reaches 45.1% next-token accuracy on easy splits versus 41.6% for the same-size distilled baseline, and gives a stronger starting point for later RL. Points to RL on unlabeled text rather than only curated verifiable tasks.

Date
Monday, 9 June 2025
Lab
Microsoft Research / Peking / Tsinghua
Kind
paper
Access
paper only

Figures

MeasureValueMeasured by
Next-token accuracy, easy split (14B)45.11% vs 41.60%
RPT-14B vs R1-Distill-Qwen-14B
authors

Preliminary, small-scale; no frontier lab has confirmed adopting it.

Sources

  1. arxiv.org/abs/2506.08007
  2. arxiv.org/html/2506.08007

This record was checked against its sources on 6 October 2026. How we check

Related