AI Research Atlas

DeepSeek-GRM / SPCT (inference-time scaling for reward models)

DeepSeek · 3 April 2025

Self-Principled Critique Tuning trains generative reward models to write their own principles and critiques, so reward quality scales with extra inference compute.

Online RL trains a pointwise generative reward model (GRM-27B) that samples principles and critiques in parallel and uses a meta reward model to vote. Aimed at rewarding non-verifiable tasks, where R1's rule-based rewards do not apply.

Date
Thursday, 3 April 2025
Lab
DeepSeek
Kind
paper
Access
open weights

arXiv 2504.02495 v1 was submitted 2025-04-03 (v3 2025-09-25), first author Zijun Liu with 7 others. Inference-time scaling here means sampling multiple principle-and-critique rollouts and voting with a meta reward model. Whether a deployed DeepSeek model used it was not found.

Sources

  1. arxiv.org/abs/2504.02495

This record was checked against its sources on 6 October 2026. How we check

Related