DeepSeek-V2 (Multi-head Latent Attention)
Introduces Multi-head Latent Attention (MLA), compressing the KV cache by 93.3% in a 236B MoE with 21B active parameters.
MLA stores a low-rank latent instead of per-head keys and values, slashing inference memory. Combined with DeepSeekMoE it cut training cost 42.5% versus DeepSeek 67B and raised generation throughput 5.76x. MLA became the backbone of V3, R1 and Kimi K2.
- Date
- Tuesday, 7 May 2024
- Lab
- DeepSeek
- Kind
- paper
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| KV cache reduction | 93.3% vs DeepSeek 67B | company |
| Training cost reduction | 42.5% vs DeepSeek 67B; 8.1T tokens | company |
| Generation throughput | 5.76x | company |
arXiv v1 2024-05-07, final version 2024-06-19. 128K context; trained on 8.1T tokens. Authored by 150+ researchers from DeepSeek-AI.
Sources
This record was checked against its sources on 6 October 2026. How we check