DeepSeek-V3 Technical Report
671B MoE (37B active) trained on 14.8T tokens with 2.788M H800 GPU-hours, no loss spikes or rollbacks; auxiliary-loss-free balancing and multi-token prediction.
Joins MLA, DeepSeekMoE, FP8 mixed-precision training, DualPipe overlap and a load-balancing scheme without auxiliary loss. The paper made frontier-class efficiency public and set up R1; the often-quoted $5.6M figure covers only the final training run's GPU rental.
- Date
- Friday, 27 December 2024
- Lab
- DeepSeek
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Parameters (total / active) | 671B / 37B 14.8T pretraining tokens | company |
| Training compute | 2.788M H800 GPU-hours final run only; excludes prior research, ablations, data | company |
arXiv v1 2024-12-27, v2 2025-02-18. Weights released 2024-12-26 with the announcement. Cost figure is company-reported and covers GPU time for the final run.
Sources
This record was checked against its sources on 6 October 2026. How we check