Native Sparse Attention (NSA)
Hardware-aligned sparse attention that is trained natively, combining token compression, selection and sliding windows, matching full attention at 64K with big speedups.
Unlike post-hoc sparsification, NSA is part of pretraining, so the model learns where to look. Gives substantial speedups over full attention in decoding, forward and backward at 64K. Led to DeepSeek Sparse Attention in V3.2.
- Date
- Sunday, 16 February 2025
- Lab
- DeepSeek
- Kind
- paper
- Access
- paper only
arXiv v1 2025-02-16. Lead author Jingyang Yuan; Liang Wenfeng is a co-author per the paper (not shown in my fetch). The ACL 2025 best-paper claim from memory was NOT confirmed on the arXiv page and is omitted.
Sources
This record was checked against its sources on 6 October 2026. How we check