Qwen2.5 (0.5B-72B) with Qwen2.5-Coder and Qwen2.5-Math
Qwen2.5 family pretrained on up to 18T tokens, sizes 0.5B to 72B, 128K context; 72B reported ahead of Llama-3.1-70B and Mistral-Large-V2.
Pretraining data rose to up to 18T tokens; better instruction following, structured data and JSON output, 29+ languages, 128K input with 8K generation. Released with Coder and Math specialists the same day. Apache 2.0 except the 3B and 72B models.
- Date
- Thursday, 19 September 2024
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Pretraining tokens | up to 18T Context 128K in, 8K out; Qwen2.5-Plus API model competitive with Llama-3.1-405B per company | company |
Qwen2.5-Coder 32B followed on 2024-11-12 (blog 'Qwen2.5-Coder Series: Powerful, Diverse, Practical'); Qwen2.5-Turbo (1M context, API) on 2024-11-15.
Sources
This record was checked against its sources on 6 October 2026. How we check