DeepSeek-R1: Incentivizing Reasoning via RL (arXiv)
Shows reasoning can emerge from pure RL on verifiable rewards (R1-Zero), then builds R1 with a small cold-start and distills it to small models.
A reproducible recipe, in which GRPO plus rule-based accuracy/format rewards with no human reasoning traces gives self-reflection, longer chains and 'aha' moments. R1 adds cold-start SFT, rejection sampling and a second RL stage. Weights released MIT-licensed, which triggered the open reasoning replication wave.
- Date
- Wednesday, 22 January 2025
- Lab
- DeepSeek
- Kind
- paper
- Access
- open weights
- Price
- $0.55 in / $2.19 out per M tokens (launch, 2025-01-20)
Figures
| Measure | Value | Measured by |
|---|---|---|
| API price (R1 launch) | $0.55 in / $2.19 out per M tokens cache-miss input; DeepSeek API docs, 2025-01-20 | company |
arXiv v1 2025-01-22; weights and API release 2025-01-20 per DeepSeek's docs. arXiv v2 posted 2026-01-04 (expanded version). Benchmarks in the abstract are company-measured; open reproductions (Open-R1 etc.) followed within weeks.
Sources
This record was checked against its sources on 6 October 2026. How we check