Qwen2.5-VL (3B, 7B, 32B, 72B)
Qwen2.5-VL: visual agent that operates computers and phones, understands 1-hour video, and emits structured JSON; 3B, 7B, 72B open.
Flagship vision-language model with object localization via boxes and points, structured output for invoices and forms, and event pinpointing in videos over an hour. 3B, 7B and 72B base and instruct models opened; a 32B variant followed 2025-03-24. Technical report arXiv 2502.13923.
- Date
- Sunday, 26 January 2025
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Technical report | arXiv 2502.13923 Submitted 2025-02-19 | company |
Blog dated 2025-01-26 (Hugging Face repos 2025-01-26 and 27); the Qwen3-VL README news lists the series on 2025-01-28. Licenses are per Wikipedia. Benchmarks not reproduced here.
Sources
- qwenlm.github.io/blog/qwen2.5-vl/
- arxiv.org/abs/2502.13923
- github.com/QwenLM/Qwen3-VL
- en.wikipedia.org/wiki/Qwen
This record was checked against its sources on 6 October 2026. How we check