Qwen3-Max-Thinking
Alibaba's trillion-parameter reasoning flagship with adaptive tool use; company claims parity with GPT-5.2-Thinking, Artificial Analysis scores it 40 on its Intelligence Index.
Adds large-scale RL and a new test-time scaling scheme with adaptive tool use on the 1T-plus, 36T-token Qwen3-Max base. Closed API with 256K context. Independent testing placed it behind DeepSeek V3.2 (42), GLM-4.7 (42) and Kimi K2.5 (47).
- Date
- Monday, 26 January 2026
- Lab
- Alibaba (Qwen)
- Kind
- model
- Access
- closed API
- Price
- $1.2 / $6 per M tokens (up to 32K input), per Artificial Analysis 2026-01
Figures
| Measure | Value | Measured by |
|---|---|---|
| Artificial Analysis Intelligence Index | 40 Behind DeepSeek V3.2 and GLM-4.7 at 42 and Kimi K2.5 at 47 | independent |
| Humanity's Last Exam | 26% (AA) vs 58.3 company with tools Alibaba's tool-assisted figure (technews.tw) is not comparable to AA's no-tools 26%; Gigazine lists 30.2 | company |
| API price per M tokens | $1.2 in / $6 out Up to 32K input; $3/$15 at 128K-256K | independent |
Announcement dated 2026-01-26 by technews.tw; API id qwen3-max-2026-01-23 per Gigazine. Alibaba benchmark claims reach me via press coverage, not the primary blog. The same coverage says Qwen passed 200,000 derivative models and 1 billion downloads (company figure).
Sources
- technews.tw/2026/01/27/qwen3-max-thinking-claims-performance-on-par-with-gpt-5/
- gigazine.net/gsc_news/en/20260127-qwen3-max-thinking/
- artificialanalysis.ai/articles/qwen3-max-thinking-everything-you-need-to-know
This record was checked against its sources on 6 October 2026. How we check