Qwen3.7-Plus
Proprietary multimodal agent model at $0.40 input and $1.60 output per million tokens, about 60% below Qwen3.7-Max; 1M context.
Unifies vision and language as a single agent foundation, adding image and video input at a fraction of Max's price. Cloud-API only. Reported ahead of DeepSeek-V4-Pro on terminal tasks and ahead of GPT-5.4 and Claude Opus 4.6 on GUI tasks.
- Date
- Sunday, 31 May 2026
- Lab
- Alibaba (Qwen)
- Kind
- model
- Access
- closed API
- Price
- $0.40 input / $1.60 output per M tokens; cached input $0.04 (2026-06)
Figures
| Measure | Value | Measured by |
|---|---|---|
| Terminal Bench 2.0-Terminus | 70.3 ScreenSpot Pro 79.0; 1M context with 256K reserved for reasoning; cached input $0.04/M | company |
Qwen blog dated 2026-05-31 (Releasebot); VentureBeat gives 2026-06-02; Model Studio lists qwen3.7-plus-2026-05-26 on 2026-06-01. Comparisons come from VentureBeat's report of Alibaba's figures. Qwen3.7-Flash followed on 2026-07-21.
Sources
- venturebeat.com/technology/alibabas-qwen3-7-plus-supports-text-video-and-imagery-inputs-at
- releasebot.io/updates/qwen
- www.alibabacloud.com/help/en/model-studio/newly-released-models
This record was checked against its sources on 6 October 2026. How we check