Qwen2.5-Turbo (1M-token context)
API model with 1M-token context that gets 100% passkey retrieval and RULER 93.1 versus GPT-4's 91.6, with a 4.3x faster first token via sparse attention.
Context raised from 128K to 1M tokens (about 10 novels), with sparse attention cutting time to first token for 1M tokens from 4.9 minutes to 68 seconds. Price unchanged at 0.3 yuan per million tokens. The open 7B and 14B Qwen2.5-1M models followed on 2025-01-27.
- Date
- Friday, 15 November 2024
- Lab
- Alibaba (Qwen)
- Kind
- model
- Access
- closed API
- Price
- 0.3 yuan per M tokens (Alibaba, 2024-11)
Figures
| Measure | Value | Measured by |
|---|---|---|
| RULER (1M context) | 93.1 vs GPT-4 91.6 and GLM4-9B-1M 89.9; Passkey 100% | company |
| Time to first token at 1M | 68 s (from 4.9 min) 4.3x speedup from sparse attention; 0.3 yuan per M tokens | company |
Scores are Alibaba-run. Blog dated 2024-11-15.
Sources
This record was checked against its sources on 6 October 2026. How we check