Tulu 3 (names RLVR)
Fully open post-training recipe (SFT, DPO, new RLVR stage) that beats Llama 3.1 Instruct and coins 'Reinforcement Learning with Verifiable Rewards'.
Replaces the learned reward model with a programmatic check (math answer, instruction constraint) for a final RL stage. Releases data, code, evals and checkpoints. The name RLVR stuck, and the same idea powered R1-Zero two months later.
- Date
- Friday, 22 November 2024
- Lab
- Allen Institute for AI (Ai2)
- Kind
- paper
- Access
- open weights
Lead author Nathan Lambert. Verifiable-reward RL had precedents (e.g. rule-based rewards in math/code RL, STaR-style filtering); Tulu 3 is the first widely read paper to name and open-source it as a post-training stage. arXiv v1 2024-11-22.
Sources
This record was checked against its sources on 6 October 2026. How we check