Attention Residuals (AttnRes)
Replaces fixed residual accumulation with softmax attention over earlier layers' outputs, fixing PreNorm dilution; Block AttnRes keeps memory manageable.
Each layer learns input-dependent weights over previous layer (or block) outputs instead of adding them uniformly. Gains across tasks in scaling-law runs and in Kimi Linear (48B/3B active, 1.4T tokens). Goes after a component unchanged since ResNets.
- Date
- Monday, 16 March 2026
- Lab
- Moonshot AI (Kimi Team)
- Kind
- paper
- Access
- paper only
arXiv v1 2026-03-16 (technical report dated 2026-03-15 per press). 36 authors from the Kimi Team (Guangyu Chen, Yu Zhang, Jianlin Su among leads).
Sources
This record was checked against its sources on 6 October 2026. How we check