Qwen2.5-Math-PRM and ProcessBench
Qwen releases open process reward models (7B, 72B) and ProcessBench, a 3,400-case benchmark for finding the first erroneous step in math reasoning.
Argues PRMs should be judged on step-level error identification as well as best-of-N, and open-sources a leading PRM plus the benchmark for scalable oversight of reasoning.
- Date
- Tuesday, 14 January 2025
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| ProcessBench size | 3,400 test cases Step-level error identification in mathematical reasoning | company |
Blog dated 2025-01-14 in the feed (HF repos 2025-01-13); released a week before DeepSeek-R1 and relevant to the outcome-vs-process reward debate.
Sources
- qwenlm.github.io/blog/qwen2.5-math-prm/
- huggingface.co/api/models?author=Qwen&sort=createdAt&direction=1&limit=100
This record was checked against its sources on 6 October 2026. How we check