Qwen3-VL (235B-A22B first; 2B to 32B and 30B-A3B in October)
Open 235B-A22B vision-language flagship with 256K context (1M extendable), OCR in 32 languages and GUI-agent skills. Smaller sizes followed in October.
Early joint text-vision pretraining keeps text ability level with Qwen3-235B-A22B-2507. Adds 3D grounding, image-to-code, GUI control and hours-long video with second-level retrieval. Instruct and Thinking variants, Apache 2.0.
- Date
- Tuesday, 23 September 2025
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Context length | 256K native, up to 1M All sizes; OCR languages rose from 10 to 32 | company |
| Technical report | arXiv 2511.21631 Submitted 2025-11-26 | company |
README dates are 235B 2025-09-23; 30B-A3B 2025-10-04; 4B/8B 2025-10-15; 2B/32B 2025-10-21. Alibaba says the Instruct model matches or beats Gemini 2.5 Pro on major visual-perception benchmarks (company claim, not independently measured).
Sources
- qwen.ai/blog?id=qwen3-vl
- github.com/QwenLM/Qwen3-VL
- huggingface.co/Qwen/Qwen3-VL-235B-A22B-Instruct
- arxiv.org/abs/2511.21631
This record was checked against its sources on 6 October 2026. How we check