AI Research Atlas

Let's Verify Step by Step (PRM800K)

OpenAI · 31 May 2023

Step-level feedback beats outcome-only feedback for training verifiers; best reward model solves 78% of a MATH subset; 800K step labels released.

Compares process reward models (labels on each reasoning step) with outcome reward models for best-of-N selection on competition math, and releases the PRM800K dataset. A direct line to later verifier and reward-model work for reasoning (inference).

Date
Wednesday, 31 May 2023
Lab
OpenAI
Kind
paper
Access
research preview

Figures

MeasureValueMeasured by
MATH test subset solved with best process reward model78%
representative subset of MATH test; best-of-N re-ranking
company

Authors include Hunter Lightman, Vineet Kosaraju, Karl Cobbe, John Schulman, Ilya Sutskever and Jan Leike. Paper-only release; the dataset PRM800K was published with it.

Sources

  1. arxiv.org/abs/2305.20050

This record was checked against its sources on 6 October 2026. How we check