FlashQLA (linear-attention kernels)
Open TileLang kernel library for Gated DeltaNet prefill, claimed 2-3x faster forward and 2x faster backward than the FLA Triton kernels on Hopper and Blackwell.
Fused, warp-specialized kernels with automatic intra-card context parallelism, aimed at training and serving the Gated DeltaNet hybrids used from Qwen3-Next onward. It later gained SM100 and SM120 support and became a backend for flash-linear-attention.
- Date
- Friday, 24 April 2026
- Lab
- Alibaba (Qwen)
- Kind
- infra
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Forward / backward speedup vs FLA Triton | 2-3x / 2x Across several scenarios on NVIDIA Hopper and Blackwell, per the README | company |
Date is GitHub repo creation (2026-04-24); the README's dated news starts at v0.1.1 in 2026-06. Software, not model weights. Speedups are Qwen's own.
Sources
This record was checked against its sources on 6 October 2026. How we check