Qwen3.8-Flash-Next
125B-parameter (6B active) multimodal MoE that previews the Qwen4 architecture, built on Gated DeltaNet plus Qwen Sparse Attention, with 51B of N-gram embeddings.
Hybrid of Gated DeltaNet and global attention (1 layer in 4), Qwen Sparse Attention at continued pretraining, a four-branch Gated Residual and 51B n-gram embeddings held in host memory. Per its design paper it leads the 397B-A17B Qwen3.5 on 8 of 14 pretraining benchmarks at about 1/3 the activated parameters and 1/9 the training FLOPs.
- Date
- Wednesday, 26 August 2026
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| SWE-bench Pro / GPQA Diamond | 62.5 / 91.7 LiveCodeBench 91.9, Toolathlon 73.5, AndroidWorld 84.5; 125B total, 6B active, 262K native context | company |
| Pretraining benchmarks vs Qwen3.5-397B-A17B | leads on 8 of 14 Trails by at most 2.6 points on the rest; about 1/3 activated parameters, 1/3 training tokens, roughly 1/9 training FLOPs | company |
Model card numbers (SWE-bench Pro 62.5, GPQA Diamond 91.7, LiveCodeBench v6 91.9, Toolathlon 73.5, AndroidWorld 84.5; 125B total, 6B active plus 51B n-gram embeddings) match the Hugging Face table. Model Studio lists the hosted qwen3.8-flash on 2026-08-26. Qwen's post calls it an early preview of the Qwen4 architecture; the n-gram embeddings resemble DeepSeek's Engram idea but influence was not checked (Inference). License is qwen-community-1.0, not Apache 2.0. Design paper 'On the Design of Qwen3.8-Next Architecture' (arXiv 2608.30320, 36 authors) was submitted 2026-08-31; its efficiency ratios are company measurements. Design paper 'On the Design of Qwen3.8-Next Architecture' (arXiv 2608.30320, 36 authors) was submitted 2026-08-31; its efficiency ratios are company measurements. Design paper 'On the Design of Qwen3.8-Next Architecture' (arXiv 2608.30320, 36 authors) was submitted 2026-08-31; its efficiency ratios are company measurements.
Sources
- huggingface.co/Qwen/Qwen3.8-Flash-Next
- releasebot.io/updates/qwen
- api.github.com/orgs/QwenLM/repos?sort=created&direction=desc&per_page=50
- www.alibabacloud.com/help/en/model-studio/newly-released-models
- arxiv.org/abs/2608.30320
This record was checked against its sources on 6 October 2026. How we check