AI Research Atlas

Qwen2.5-1M (open models with 1M-token context)

Alibaba (Qwen) · 27 January 2025

Open Qwen2.5-7B and 14B Instruct-1M with a vLLM-based framework that processes 1M-token inputs 3x to 7x faster.

First time Qwen upgraded its open models to 1M-token contexts, two months after the API-only Qwen2.5-Turbo. Adds sparse-attention inference support and a technical report on training and inference design.

Date
Monday, 27 January 2025
Lab
Alibaba (Qwen)
Kind
open-weights
Access
open weights

Figures

MeasureValueMeasured by
Inference speedup at 1M tokens3x to 7x
Open-sourced vLLM-based framework with sparse attention (company)
company

Blog dated 2025-01-27 in the feed (Hugging Face repos created 2025-01-23).

Sources

  1. qwenlm.github.io/blog/qwen2.5-1m/
  2. huggingface.co/api/models?author=Qwen&sort=createdAt&direction=1&limit=100

This record was checked against its sources on 6 October 2026. How we check

Related