AI Research Atlas

DeepSeek-V3

DeepSeek · 26 December 2024

671B MoE (37B active) trained on 14.8T tokens in 2.788M H800 GPU hours; open weights matching leading closed models at far lower reported cost.

MLA plus DeepSeekMoE with auxiliary-loss-free load balancing, multi-token prediction and FP8 mixed-precision training, with no irrecoverable loss spikes. Post-trained with distillation from an R1-style reasoner. API ran at $0.27/$1.10 per million tokens after a promo.

Date
Thursday, 26 December 2024
Lab
DeepSeek
Kind
open-weights
Access
open weights (restricted license)
Price
$0.27 input (miss) / $0.07 (hit) / $1.10 output per M tokens from 2025-02-08; V2 rates until then

Figures

MeasureValueMeasured by
Pretraining compute2.788M H800 GPU-hours
14.8T tokens; the widely quoted ~$5.6M covers only the final run at assumed $2/GPU-hour
company
MMLU / MMLU-Pro87.1 / 64.4
Base model; GSM8K 89.3, HumanEval 65.2
company

The $5.58M figure is final-run rental-equivalent compute only; it excludes prior research, ablations, data, salaries and the ~$51M+ hardware purchase estimated by The Register. Weights repo created 2024-12-25 UTC; announcement 2024-12-26.

Sources

  1. arxiv.org/abs/2412.19437
  2. api-docs.deepseek.com/news/news1226
  3. huggingface.co/deepseek-ai/DeepSeek-V3
  4. www.theregister.com/2025/09/19/deepseek_cost_train/

This record was checked against its sources on 6 October 2026. How we check

Related