AI Research Atlas

DeepSeek-R1 and R1-Zero

DeepSeek · 20 January 2025

Open-weights 671B MoE reasoning model claimed at OpenAI o1 level; R1-Zero showed reasoning emerging from pure RL with rule-based rewards.

GRPO with verifiable (rule-based) rewards on a V3 base; R1-Zero used no SFT and showed self-reflection emerging. R1 adds a cold-start SFT and multi-stage RL. MIT license, outputs allowed for distillation, six distilled 1.5B-70B models.

Date
Monday, 20 January 2025
Lab
DeepSeek
Kind
open-weights
Access
open weights
Price
$0.14 (hit) / $0.55 (miss) input, $2.19 output per M tokens (2025-01-20)

Figures

MeasureValueMeasured by
AIME 2024 pass@179.8%
MATH-500 97.3%, GPQA Diamond 71.5%, Codeforces 2029, LiveCodeBench 65.9%
company

Scores are DeepSeek-reported. The arXiv paper (2501.12948) was posted 2025-01-22 and later published in Nature (2025-09-17). Distilled models are Qwen2.5 and Llama3 fine-tunes, not R1 itself.

Sources

  1. arxiv.org/abs/2501.12948
  2. api-docs.deepseek.com/news/news250120
  3. huggingface.co/deepseek-ai/DeepSeek-R1

This record was checked against its sources on 6 October 2026. How we check

Related