Native Sparse Attention (NSA)
Trainable, hardware-aligned sparse attention that combines coarse token compression with fine-grained selection and matches or beats full attention at 64K context.
Unlike inference-only sparsity, NSA is natively trainable end to end and designed around GPU arithmetic intensity, giving large decode, forward and backward speedups at 64K length. Inferred to be a conceptual precursor to DeepSeek Sparse Attention in V3.2.
- Date
- Sunday, 16 February 2025
- Lab
- DeepSeek
- Kind
- paper
- Access
- research preview
Paper only; the production successor, DSA, shipped in V3.2-Exp (2025-09-29). Whether DSA descends directly from NSA is an inference from shared design goals; the V3.2 report should be checked for explicit attribution.
Sources
This record was checked against its sources on 6 October 2026. How we check