Kimi Linear
Hybrid linear-attention architecture (Kimi Delta Attention + MLA) that beats full attention in fair comparisons while cutting KV cache up to 75%.
KDA extends Gated DeltaNet with finer-grained gating. In a 48B-total/3B-active model it outperforms full-MLA with the same recipe and delivers up to 6x decoding throughput at 1M context. KDA reappears in Kimi K3.
- Date
- Thursday, 30 October 2025
- Lab
- Moonshot AI
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| KV cache reduction | up to 75% vs full MLA | company |
| Decoding throughput at 1M context | up to 6x vs full MLA | company |
MIT license. Claim of outperforming full attention is Moonshot's, at the 3B-active scale tested.
Sources
- arxiv.org/abs/2510.26692
- huggingface.co/moonshotai/Kimi-Linear-48B-A3B-Instruct
- www.kimi.com/blog/kimi-k3
This record was checked against its sources on 6 October 2026. How we check