DeepSeek-R1 and R1-Zero
Open-weights 671B MoE reasoning model claimed at OpenAI o1 level; R1-Zero showed reasoning emerging from pure RL with rule-based rewards.
GRPO with verifiable (rule-based) rewards on a V3 base; R1-Zero used no SFT and showed self-reflection emerging. R1 adds a cold-start SFT and multi-stage RL. MIT license, outputs allowed for distillation, six distilled 1.5B-70B models.
- Date
- Monday, 20 January 2025
- Lab
- DeepSeek
- Kind
- open-weights
- Access
- open weights
- Price
- $0.14 (hit) / $0.55 (miss) input, $2.19 output per M tokens (2025-01-20)
Figures
| Measure | Value | Measured by |
|---|---|---|
| AIME 2024 pass@1 | 79.8% MATH-500 97.3%, GPQA Diamond 71.5%, Codeforces 2029, LiveCodeBench 65.9% | company |
Scores are DeepSeek-reported. The arXiv paper (2501.12948) was posted 2025-01-22 and later published in Nature (2025-09-17). Distilled models are Qwen2.5 and Llama3 fine-tunes, not R1 itself.
Sources
This record was checked against its sources on 6 October 2026. How we check