AI Research Atlas

DeepSeek-V3 Technical Report

DeepSeek · 27 December 2024

671B MoE (37B active) trained on 14.8T tokens with 2.788M H800 GPU-hours, no loss spikes or rollbacks; auxiliary-loss-free balancing and multi-token prediction.

Joins MLA, DeepSeekMoE, FP8 mixed-precision training, DualPipe overlap and a load-balancing scheme without auxiliary loss. The paper made frontier-class efficiency public and set up R1; the often-quoted $5.6M figure covers only the final training run's GPU rental.

Date
Friday, 27 December 2024
Lab
DeepSeek
Kind
paper
Access
open weights

Figures

MeasureValueMeasured by
Parameters (total / active)671B / 37B
14.8T pretraining tokens
company
Training compute2.788M H800 GPU-hours
final run only; excludes prior research, ablations, data
company

arXiv v1 2024-12-27, v2 2025-02-18. Weights released 2024-12-26 with the announcement. Cost figure is company-reported and covers GPU time for the final run.

Sources

  1. arxiv.org/abs/2412.19437

This record was checked against its sources on 6 October 2026. How we check

Related