AI Research Atlas

Kimi Linear

Moonshot AI · 30 October 2025

Hybrid linear-attention architecture (Kimi Delta Attention + MLA) that beats full attention in fair comparisons while cutting KV cache up to 75%.

KDA extends Gated DeltaNet with finer-grained gating. In a 48B-total/3B-active model it outperforms full-MLA with the same recipe and delivers up to 6x decoding throughput at 1M context. KDA reappears in Kimi K3.

Date
Thursday, 30 October 2025
Lab
Moonshot AI
Kind
open-weights
Access
open weights

Figures

MeasureValueMeasured by
KV cache reductionup to 75%
vs full MLA
company
Decoding throughput at 1M contextup to 6x
vs full MLA
company

MIT license. Claim of outperforming full attention is Moonshot's, at the 3B-active scale tested.

Sources

  1. arxiv.org/abs/2510.26692
  2. huggingface.co/moonshotai/Kimi-Linear-48B-A3B-Instruct
  3. www.kimi.com/blog/kimi-k3

This record was checked against its sources on 6 October 2026. How we check

Related