AI Research Atlas

Qwen2 (0.5B, 1.5B, 7B, 57B-A14B MoE, 72B)

Alibaba (Qwen) · 7 June 2024

Five Qwen2 sizes including a 57B-A14B MoE and a 72B, 27 added languages, grouped-query attention everywhere and 128K context in the 7B and 72B.

Dense 0.5B, 1.5B, 7B and 72B plus a 57B-A14B MoE (57.41B total). All sizes use GQA; 7B and 72B instruct models reach 128K context (57B-A14B 64K); data covers 27 additional languages. Per Wikipedia, the 72B keeps the Tongyi Qianwen license and the rest are Apache 2.0.

Date
Friday, 7 June 2024
Lab
Alibaba (Qwen)
Kind
open-weights
Access
open weights

Figures

MeasureValueMeasured by
Context length128K (7B, 72B), 64K (57B-A14B), 32K (0.5B, 1.5B)
Parameters 0.49B to 72.71B; 27 extra languages beyond English and Chinese
company

Blog 'Hello Qwen2' dated 2024-06-07; some weights existed on HF from 2024-05-22. Licenses are per Wikipedia, not the blog. Benchmark tables were not opened.

Sources

  1. qwenlm.github.io/blog/qwen2/
  2. en.wikipedia.org/wiki/Qwen
  3. huggingface.co/api/models?author=Qwen&sort=createdAt&direction=1&limit=100

This record was checked against its sources on 6 October 2026. How we check

Related