AI Research Atlas

DeepSeek-V2 (Multi-head Latent Attention)

DeepSeek · 7 May 2024

Introduces Multi-head Latent Attention (MLA), compressing the KV cache by 93.3% in a 236B MoE with 21B active parameters.

MLA stores a low-rank latent instead of per-head keys and values, slashing inference memory. Combined with DeepSeekMoE it cut training cost 42.5% versus DeepSeek 67B and raised generation throughput 5.76x. MLA became the backbone of V3, R1 and Kimi K2.

Date
Tuesday, 7 May 2024
Lab
DeepSeek
Kind
paper
Access
open weights

Figures

MeasureValueMeasured by
KV cache reduction93.3%
vs DeepSeek 67B
company
Training cost reduction42.5%
vs DeepSeek 67B; 8.1T tokens
company
Generation throughput5.76xcompany

arXiv v1 2024-05-07, final version 2024-06-19. 128K context; trained on 8.1T tokens. Authored by 150+ researchers from DeepSeek-AI.

Sources

  1. arxiv.org/abs/2405.04434

This record was checked against its sources on 6 October 2026. How we check

Related