Qwen2 (0.5B, 1.5B, 7B, 57B-A14B MoE, 72B)
Five Qwen2 sizes including a 57B-A14B MoE and a 72B, 27 added languages, grouped-query attention everywhere and 128K context in the 7B and 72B.
Dense 0.5B, 1.5B, 7B and 72B plus a 57B-A14B MoE (57.41B total). All sizes use GQA; 7B and 72B instruct models reach 128K context (57B-A14B 64K); data covers 27 additional languages. Per Wikipedia, the 72B keeps the Tongyi Qianwen license and the rest are Apache 2.0.
- Date
- Friday, 7 June 2024
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Context length | 128K (7B, 72B), 64K (57B-A14B), 32K (0.5B, 1.5B) Parameters 0.49B to 72.71B; 27 extra languages beyond English and Chinese | company |
Blog 'Hello Qwen2' dated 2024-06-07; some weights existed on HF from 2024-05-22. Licenses are per Wikipedia, not the blog. Benchmark tables were not opened.
Sources
- qwenlm.github.io/blog/qwen2/
- en.wikipedia.org/wiki/Qwen
- huggingface.co/api/models?author=Qwen&sort=createdAt&direction=1&limit=100
This record was checked against its sources on 6 October 2026. How we check