Qwen3-Next-80B-A3B (Instruct and Thinking)
80B MoE with 3B active, mixing Gated DeltaNet and gated attention 3:1; claimed 10x throughput beyond 32K context at 10% of Qwen3-32B training cost.
Ultra-sparse MoE (512 experts, 10 routed plus 1 shared), multi-token prediction and stability tweaks such as zero-centered, weight-decayed layernorm. Instruct matches Qwen3-235B-A22B-Instruct-2507 on some benchmarks and is stronger at 256K context. The hybrid layout was reused in Qwen3-Coder-Next and Qwen3.5.
- Date
- Thursday, 11 September 2025
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Base model training cost vs Qwen3-32B-Base | about 10% While outperforming it on downstream tasks, per Qwen | company |
| Inference throughput vs Qwen3-32B-Base | more than 10x For context over 32K tokens | company |
| Pretraining tokens | 15T 80B total parameters, 3B active; context 262,144 native, about 1M extended | company |
GitHub README dates the release 2025-09-11 (blog timestamp 2025-09-10 UTC; Hugging Face repos created 2025-09-09). Throughput and cost multiples are company claims. Apache 2.0.
Sources
- huggingface.co/Qwen/Qwen3-Next-80B-A3B-Instruct
- huggingface.co/Qwen/Qwen3-Next-80B-A3B-Thinking
- github.com/QwenLM/Qwen3.8
- the-decoder.com/alibabas-qwen3-next-builds-on-a-faster-moe-architecture/
This record was checked against its sources on 6 October 2026. How we check