AI Research Atlas

Tulu 3 (names RLVR)

Allen Institute for AI (Ai2) · 22 November 2024

Fully open post-training recipe (SFT, DPO, new RLVR stage) that beats Llama 3.1 Instruct and coins 'Reinforcement Learning with Verifiable Rewards'.

Replaces the learned reward model with a programmatic check (math answer, instruction constraint) for a final RL stage. Releases data, code, evals and checkpoints. The name RLVR stuck, and the same idea powered R1-Zero two months later.

Date
Friday, 22 November 2024
Lab
Allen Institute for AI (Ai2)
Kind
paper
Access
open weights

Lead author Nathan Lambert. Verifiable-reward RL had precedents (e.g. rule-based rewards in math/code RL, STaR-style filtering); Tulu 3 is the first widely read paper to name and open-source it as a post-training stage. arXiv v1 2024-11-22.

Sources

  1. arxiv.org/abs/2411.15124

This record was checked against its sources on 6 October 2026. How we check