Qwen3-VL-Embedding and Qwen3-VL-Reranker (2B, 8B)
Open multimodal retrieval pair built on Qwen3-VL: embeddings and rerankers for text, images, screenshots and video across 30+ languages.
Extends the text-only Qwen3 Embedding line to mixed-modality retrieval, with one embedding space for text, images, screenshots and video and a cross-encoder reranker for (query, document) pairs. Flexible vector sizes and instruction-conditioned embeddings. Apache 2.0.
- Date
- Thursday, 8 January 2026
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Technical report | arXiv 2601.04720 Submitted 2026-01-08; Matryoshka embeddings, 32K-token inputs, 2B and 8B sizes | company |
Hugging Face repos created 2026-01-07; GitHub repo created 2026-01-08. No retrieval-benchmark table was opened.
Sources
- huggingface.co/Qwen/Qwen3-VL-Embedding-8B
- github.com/QwenLM/Qwen3-VL-Embedding
- arxiv.org/abs/2601.04720
This record was checked against its sources on 6 October 2026. How we check