AI Research Atlas

Attention Residuals (AttnRes)

Moonshot AI (Kimi Team) · 16 March 2026

Replaces fixed residual accumulation with softmax attention over earlier layers' outputs, fixing PreNorm dilution; Block AttnRes keeps memory manageable.

Each layer learns input-dependent weights over previous layer (or block) outputs instead of adding them uniformly. Gains across tasks in scaling-law runs and in Kimi Linear (48B/3B active, 1.4T tokens). Goes after a component unchanged since ResNets.

Date
Monday, 16 March 2026
Lab
Moonshot AI (Kimi Team)
Kind
paper
Access
paper only

arXiv v1 2026-03-16 (technical report dated 2026-03-15 per press). 36 authors from the Kimi Team (Guangyu Chen, Yu Zhang, Jianlin Su among leads).

Sources

  1. arxiv.org/abs/2603.15031

This record was checked against its sources on 6 October 2026. How we check

Related