AI Research Atlas

Qwen2.5-Turbo (1M-token context)

Alibaba (Qwen) · 15 November 2024

API model with 1M-token context that gets 100% passkey retrieval and RULER 93.1 versus GPT-4's 91.6, with a 4.3x faster first token via sparse attention.

Context raised from 128K to 1M tokens (about 10 novels), with sparse attention cutting time to first token for 1M tokens from 4.9 minutes to 68 seconds. Price unchanged at 0.3 yuan per million tokens. The open 7B and 14B Qwen2.5-1M models followed on 2025-01-27.

Date
Friday, 15 November 2024
Lab
Alibaba (Qwen)
Kind
model
Access
closed API
Price
0.3 yuan per M tokens (Alibaba, 2024-11)

Figures

MeasureValueMeasured by
RULER (1M context)93.1
vs GPT-4 91.6 and GLM4-9B-1M 89.9; Passkey 100%
company
Time to first token at 1M68 s (from 4.9 min)
4.3x speedup from sparse attention; 0.3 yuan per M tokens
company

Scores are Alibaba-run. Blog dated 2024-11-15.

Sources

  1. qwenlm.github.io/blog/qwen2.5-turbo/

This record was checked against its sources on 6 October 2026. How we check

Related