Qwen1.5-MoE-A2.7B
First Qwen MoE: 2.7B activated parameters match 7B dense models with a claimed 75% lower training cost and 1.74x faster inference.
Fine-grained experts, 'upcycling' initialization and a mix of shared and routed experts; the blog cites DeepSeek-MoE and DBRX as precedent for fine-grained experts. Stepping stone to Qwen2-57B-A14B and Qwen3's MoE line.
- Date
- Thursday, 28 March 2024
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights (restricted license)
Figures
| Measure | Value | Measured by |
|---|---|---|
| Training cost vs Qwen1.5-7B | -75% Alibaba claim; 2.7B activated, 2.0B non-embedding parameters vs 6.5B for Qwen1.5-7B | company |
| Inference speed vs Qwen1.5-7B | 1.74x Company-measured | company |
Blog dated 2024-03-28; HF base repo created 2024-02-29. The post names DeepSeek-MoE (arXiv 2401.06066, January 2024) as prior art for fine-grained experts, which is direct evidence of diffusion. Performance parity claims are Alibaba's.
Sources
- qwenlm.github.io/blog/qwen-moe/
- huggingface.co/api/models?author=Qwen&sort=createdAt&direction=1&limit=100
This record was checked against its sources on 6 October 2026. How we check