DeepSelect TopK kernels and DeepJIT runtime
DeepSeek opens its sparse-attention TopK kernel (2-20x faster than torch.topk) and a lightweight CUDA and Ascend JIT runtime.
DeepSelect v1.0.0 implements the TopK step used by DeepSeek Sparse Attention in V3.2, V4 and V4.1, plus samplers, on CUDA and Huawei Ascend. DeepJIT is a header-only C++20 runtime that compiles and caches kernels for both backends.
- Date
- Thursday, 10 September 2026
- Lab
- DeepSeek
- Kind
- infra
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| TopK speedup vs torch.topk | 2-20x DeepSelect v1.0.0, per its README; Ascend kernels added 2026-09-30 | company |
DeepSelect v1.0.0 and its analysis were released 2026-09-10 per the README; the DeepJIT repo was created 2026-09-08. Software, not model weights; speedups are DeepSeek's own.
Sources
This record was checked against its sources on 6 October 2026. How we check