Qwen3-235B-A22B-Instruct-2507 and Thinking-2507
Qwen3's flagship split into separate non-thinking and thinking checkpoints with 256K context; Instruct-2507 lifts AIME25 from 24.7 to 70.3.
Alibaba dropped April's single hybrid-mode model for dedicated Instruct-2507 (non-thinking only) and Thinking-2507 (thinking only) versions, each with 262,144-token native context extendable to about 1M. 30B-A3B and 4B 2507 versions followed within two weeks. Apache 2.0.
- Date
- Monday, 21 July 2025
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| AIME25 (Instruct-2507) | 70.3 vs 24.7 for the April Qwen3-235B-A22B non-thinking mode; Arena-Hard v2 79.2, GPQA 77.5 | company |
| AIME25 (Thinking-2507) | 92.3 HMMT25 83.9, LiveCodeBench v6 74.1, GPQA 81.1; context 262,144 native | company |
Hugging Face repo creation dates were Instruct 2025-07-21, Thinking 2025-07-25, 30B-A3B 2025-07-28/29 and 4B 2025-08-05. The model cards state each variant supports only one mode; Qwen's stated reason for abandoning hybrid mode was not found. All scores are Alibaba-reported.
Sources
- huggingface.co/Qwen/Qwen3-235B-A22B-Instruct-2507
- huggingface.co/Qwen/Qwen3-235B-A22B-Thinking-2507
- huggingface.co/Qwen/Qwen3-30B-A3B-Instruct-2507
- huggingface.co/Qwen/Qwen3-4B-Instruct-2507
This record was checked against its sources on 6 October 2026. How we check