AI Research Atlas

Native Sparse Attention (NSA)

DeepSeek · 16 February 2025

Trainable, hardware-aligned sparse attention that combines coarse token compression with fine-grained selection and matches or beats full attention at 64K context.

Unlike inference-only sparsity, NSA is natively trainable end to end and designed around GPU arithmetic intensity, giving large decode, forward and backward speedups at 64K length. Inferred to be a conceptual precursor to DeepSeek Sparse Attention in V3.2.

Date
Sunday, 16 February 2025
Lab
DeepSeek
Kind
paper
Access
research preview

Paper only; the production successor, DSA, shipped in V3.2-Exp (2025-09-29). Whether DSA descends directly from NSA is an inference from shared design goals; the V3.2 report should be checked for explicit attribution.

Sources

  1. arxiv.org/abs/2502.11089

This record was checked against its sources on 6 October 2026. How we check

Related