AI Research Atlas

DeepSeek-R1: Incentivizing Reasoning via RL (arXiv)

DeepSeek · 22 January 2025

Shows reasoning can emerge from pure RL on verifiable rewards (R1-Zero), then builds R1 with a small cold-start and distills it to small models.

A reproducible recipe, in which GRPO plus rule-based accuracy/format rewards with no human reasoning traces gives self-reflection, longer chains and 'aha' moments. R1 adds cold-start SFT, rejection sampling and a second RL stage. Weights released MIT-licensed, which triggered the open reasoning replication wave.

Date
Wednesday, 22 January 2025
Lab
DeepSeek
Kind
paper
Access
open weights
Price
$0.55 in / $2.19 out per M tokens (launch, 2025-01-20)

Figures

MeasureValueMeasured by
API price (R1 launch)$0.55 in / $2.19 out per M tokens
cache-miss input; DeepSeek API docs, 2025-01-20
company

arXiv v1 2025-01-22; weights and API release 2025-01-20 per DeepSeek's docs. arXiv v2 posted 2026-01-04 (expanded version). Benchmarks in the abstract are company-measured; open reproductions (Open-R1 etc.) followed within weeks.

Sources

  1. arxiv.org/abs/2501.12948
  2. api-docs.deepseek.com/news/news250120

This record was checked against its sources on 6 October 2026. How we check

Related