AI Research Atlas

Native Sparse Attention (NSA)

DeepSeek · 16 February 2025

Hardware-aligned sparse attention that is trained natively, combining token compression, selection and sliding windows, matching full attention at 64K with big speedups.

Unlike post-hoc sparsification, NSA is part of pretraining, so the model learns where to look. Gives substantial speedups over full attention in decoding, forward and backward at 64K. Led to DeepSeek Sparse Attention in V3.2.

Date
Sunday, 16 February 2025
Lab
DeepSeek
Kind
paper
Access
paper only

arXiv v1 2025-02-16. Lead author Jingyang Yuan; Liang Wenfeng is a co-author per the paper (not shown in my fetch). The ACL 2025 best-paper claim from memory was NOT confirmed on the arXiv page and is omitted.

Sources

  1. arxiv.org/abs/2502.11089

This record was checked against its sources on 6 October 2026. How we check