AI Research Atlas

Qwen2.5 (0.5B-72B) with Qwen2.5-Coder and Qwen2.5-Math

Alibaba (Qwen) · 19 September 2024

Qwen2.5 family pretrained on up to 18T tokens, sizes 0.5B to 72B, 128K context; 72B reported ahead of Llama-3.1-70B and Mistral-Large-V2.

Pretraining data rose to up to 18T tokens; better instruction following, structured data and JSON output, 29+ languages, 128K input with 8K generation. Released with Coder and Math specialists the same day. Apache 2.0 except the 3B and 72B models.

Date
Thursday, 19 September 2024
Lab
Alibaba (Qwen)
Kind
open-weights
Access
open weights

Figures

MeasureValueMeasured by
Pretraining tokensup to 18T
Context 128K in, 8K out; Qwen2.5-Plus API model competitive with Llama-3.1-405B per company
company

Qwen2.5-Coder 32B followed on 2024-11-12 (blog 'Qwen2.5-Coder Series: Powerful, Diverse, Practical'); Qwen2.5-Turbo (1M context, API) on 2024-11-15.

Sources

  1. qwenlm.github.io/blog/qwen2.5/
  2. qwenlm.github.io/blog/index.xml

This record was checked against its sources on 6 October 2026. How we check

Related