AI Research Atlas

Qwen2.5-Math-PRM and ProcessBench

Alibaba (Qwen) · 14 January 2025

Qwen releases open process reward models (7B, 72B) and ProcessBench, a 3,400-case benchmark for finding the first erroneous step in math reasoning.

Argues PRMs should be judged on step-level error identification as well as best-of-N, and open-sources a leading PRM plus the benchmark for scalable oversight of reasoning.

Date
Tuesday, 14 January 2025
Lab
Alibaba (Qwen)
Kind
open-weights
Access
open weights

Figures

MeasureValueMeasured by
ProcessBench size3,400 test cases
Step-level error identification in mathematical reasoning
company

Blog dated 2025-01-14 in the feed (HF repos 2025-01-13); released a week before DeepSeek-R1 and relevant to the outcome-vs-process reward debate.

Sources

  1. qwenlm.github.io/blog/qwen2.5-math-prm/
  2. huggingface.co/api/models?author=Qwen&sort=createdAt&direction=1&limit=100

This record was checked against its sources on 6 October 2026. How we check

Related