DeepSeek-V2
236B MoE (21B active) that introduced Multi-head Latent Attention, cutting the KV cache 93.3% and training cost 42.5% versus DeepSeek 67B.
MLA compresses keys and values into a low-rank latent, shrinking the KV cache and boosting generation throughput 5.76x; combined with DeepSeekMoE. Pretrained on 8.1T tokens, 128K context. MLA became the attention backbone of V3, R1 and later models.
- Date
- Monday, 6 May 2024
- Lab
- DeepSeek
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| KV cache reduction vs DeepSeek 67B | 93.3% Training cost -42.5%, max generation throughput 5.76x | company |
| MMLU | 78.5 V2-Lite (16B, 2.4B active) scored 58.3 | company |
V2's very low API pricing is widely credited with starting the 2024 Chinese LLM price war; I could not open a primary source for the price figures, so that claim is left out of the numbers.
Sources
This record was checked against its sources on 6 October 2026. How we check