DeepSeek-GRM / SPCT (inference-time scaling for reward models)
Self-Principled Critique Tuning trains generative reward models to write their own principles and critiques, so reward quality scales with extra inference compute.
Online RL trains a pointwise generative reward model (GRM-27B) that samples principles and critiques in parallel and uses a meta reward model to vote. Aimed at rewarding non-verifiable tasks, where R1's rule-based rewards do not apply.
- Date
- Thursday, 3 April 2025
- Lab
- DeepSeek
- Kind
- paper
- Access
- open weights
arXiv 2504.02495 v1 was submitted 2025-04-03 (v3 2025-09-25), first author Zijun Liu with 7 others. Inference-time scaling here means sampling multiple principle-and-critique rollouts and voting with a meta reward model. Whether a deployed DeepSeek model used it was not found.
Sources
This record was checked against its sources on 6 October 2026. How we check