Qwen3-Max
Alibaba's first model above 1T parameters, trained on 36T tokens, scores SWE-bench Verified 69.6 and Tau2-Bench 74.8; closed API only.
Qwen3-style MoE scaled past 1T parameters with global-batch load balancing. Alibaba reports no loss spikes in pretraining, 30% higher MFU than Qwen2.5-Max and 1M-token context training via ChunkFlow. Weights stay closed; a Thinking version was still in training.
- Date
- Wednesday, 24 September 2025
- Lab
- Alibaba (Qwen)
- Kind
- model
- Access
- closed API
Figures
| Measure | Value | Measured by |
|---|---|---|
| SWE-bench Verified | 69.6 Qwen3-Max-Instruct | company |
| Tau2-Bench | 74.8 Alibaba says this beats Claude Opus 4 and DeepSeek V3.1 | company |
| Pretraining tokens | 36T Over 1T parameters | company |
Announced at the Apsara Conference; Model Studio id qwen3-max-2025-09-23. Alibaba also claimed 100% on AIME25 and HMMT for an unreleased Thinking (heavy) variant with a code interpreter and parallel test-time compute; that variant shipped as Qwen3-Max-Thinking in January 2026.
Sources
- qwen.ai/blog?id=qwen3-max
- www.alibabacloud.com/blog/alibaba-cloud-unveils-strategic-roadmaps-for-the-next-generation
- www.alibabacloud.com/help/en/model-studio/newly-released-models
This record was checked against its sources on 6 October 2026. How we check