Qwen2.5-1M (open models with 1M-token context)
Open Qwen2.5-7B and 14B Instruct-1M with a vLLM-based framework that processes 1M-token inputs 3x to 7x faster.
First time Qwen upgraded its open models to 1M-token contexts, two months after the API-only Qwen2.5-Turbo. Adds sparse-attention inference support and a technical report on training and inference design.
- Date
- Monday, 27 January 2025
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Inference speedup at 1M tokens | 3x to 7x Open-sourced vLLM-based framework with sparse attention (company) | company |
Blog dated 2025-01-27 in the feed (Hugging Face repos created 2025-01-23).
Sources
- qwenlm.github.io/blog/qwen2.5-1m/
- huggingface.co/api/models?author=Qwen&sort=createdAt&direction=1&limit=100
This record was checked against its sources on 6 October 2026. How we check