AI Research Atlas

Qwen3-235B-A22B-Instruct-2507 and Thinking-2507

Alibaba (Qwen) · 21 July 2025

Qwen3's flagship split into separate non-thinking and thinking checkpoints with 256K context; Instruct-2507 lifts AIME25 from 24.7 to 70.3.

Alibaba dropped April's single hybrid-mode model for dedicated Instruct-2507 (non-thinking only) and Thinking-2507 (thinking only) versions, each with 262,144-token native context extendable to about 1M. 30B-A3B and 4B 2507 versions followed within two weeks. Apache 2.0.

Date
Monday, 21 July 2025
Lab
Alibaba (Qwen)
Kind
open-weights
Access
open weights

Figures

MeasureValueMeasured by
AIME25 (Instruct-2507)70.3
vs 24.7 for the April Qwen3-235B-A22B non-thinking mode; Arena-Hard v2 79.2, GPQA 77.5
company
AIME25 (Thinking-2507)92.3
HMMT25 83.9, LiveCodeBench v6 74.1, GPQA 81.1; context 262,144 native
company

Hugging Face repo creation dates were Instruct 2025-07-21, Thinking 2025-07-25, 30B-A3B 2025-07-28/29 and 4B 2025-08-05. The model cards state each variant supports only one mode; Qwen's stated reason for abandoning hybrid mode was not found. All scores are Alibaba-reported.

Sources

  1. huggingface.co/Qwen/Qwen3-235B-A22B-Instruct-2507
  2. huggingface.co/Qwen/Qwen3-235B-A22B-Thinking-2507
  3. huggingface.co/Qwen/Qwen3-30B-A3B-Instruct-2507
  4. huggingface.co/Qwen/Qwen3-4B-Instruct-2507

This record was checked against its sources on 6 October 2026. How we check

Related