DeepSeek-V3
671B MoE (37B active) trained on 14.8T tokens in 2.788M H800 GPU hours; open weights matching leading closed models at far lower reported cost.
MLA plus DeepSeekMoE with auxiliary-loss-free load balancing, multi-token prediction and FP8 mixed-precision training, with no irrecoverable loss spikes. Post-trained with distillation from an R1-style reasoner. API ran at $0.27/$1.10 per million tokens after a promo.
- Date
- Thursday, 26 December 2024
- Lab
- DeepSeek
- Kind
- open-weights
- Access
- open weights (restricted license)
- Price
- $0.27 input (miss) / $0.07 (hit) / $1.10 output per M tokens from 2025-02-08; V2 rates until then
Figures
| Measure | Value | Measured by |
|---|---|---|
| Pretraining compute | 2.788M H800 GPU-hours 14.8T tokens; the widely quoted ~$5.6M covers only the final run at assumed $2/GPU-hour | company |
| MMLU / MMLU-Pro | 87.1 / 64.4 Base model; GSM8K 89.3, HumanEval 65.2 | company |
The $5.58M figure is final-run rental-equivalent compute only; it excludes prior research, ablations, data, salaries and the ~$51M+ hardware purchase estimated by The Register. Weights repo created 2024-12-25 UTC; announcement 2024-12-26.
Sources
- arxiv.org/abs/2412.19437
- api-docs.deepseek.com/news/news1226
- huggingface.co/deepseek-ai/DeepSeek-V3
- www.theregister.com/2025/09/19/deepseek_cost_train/
This record was checked against its sources on 6 October 2026. How we check