Qwen2.5 Technical Report
The Qwen2.5 report describes pretraining data scaled from 7T to 18T tokens and post-training with 1M+ SFT samples and multistage RL.
Documents the recipe behind the Qwen2.5 family (open sizes plus hosted Turbo and Plus MoE variants, which it compares with GPT-4o-mini and GPT-4o) after the September 2024 release; v2 appeared 2025-01-03.
- Date
- Thursday, 19 December 2024
- Lab
- Alibaba (Qwen)
- Kind
- paper
- Access
- research preview
Figures
| Measure | Value | Measured by |
|---|---|---|
| Pretraining tokens | 18T (from 7T) Over 1M SFT samples and multistage RL, per the abstract | company |
arXiv 2412.15115 submitted 2024-12-19 (v2 2025-01-03). Company-authored description of its own training.
Sources
This record was checked against its sources on 6 October 2026. How we check