AI Research Atlas

DeepSelect TopK kernels and DeepJIT runtime

DeepSeek · 10 September 2026

DeepSeek opens its sparse-attention TopK kernel (2-20x faster than torch.topk) and a lightweight CUDA and Ascend JIT runtime.

DeepSelect v1.0.0 implements the TopK step used by DeepSeek Sparse Attention in V3.2, V4 and V4.1, plus samplers, on CUDA and Huawei Ascend. DeepJIT is a header-only C++20 runtime that compiles and caches kernels for both backends.

Date
Thursday, 10 September 2026
Lab
DeepSeek
Kind
infra
Access
open weights

Figures

MeasureValueMeasured by
TopK speedup vs torch.topk2-20x
DeepSelect v1.0.0, per its README; Ascend kernels added 2026-09-30
company

DeepSelect v1.0.0 and its analysis were released 2026-09-10 per the README; the DeepJIT repo was created 2026-09-08. Software, not model weights; speedups are DeepSeek's own.

Sources

  1. github.com/deepseek-ai/DeepSelect
  2. github.com/deepseek-ai/DeepJIT

This record was checked against its sources on 6 October 2026. How we check

Related