Qwen-VLA (vision-language-action generalist)
Unified vision-language-action model on Qwen3.5-4B plus a 1.15B flow-matching action decoder, covering manipulation, navigation and trajectory prediction.
One generalist trained jointly on robot manipulation, human egocentric, simulation and navigation data. Qwen says it matches or beats specialists fine-tuned per benchmark. The June Qwen-Robot Suite ships separate navigation and manipulation models; their link to Qwen-VLA was not checked.
- Date
- Thursday, 28 May 2026
- Lab
- Alibaba (Qwen)
- Kind
- model
- Access
- open weights
Figures
| Measure | Value | Measured by |
|---|---|---|
| Technical report | arXiv 2605.30280 Revised v2 on 2026-06-01 | company |
Date is GitHub repo creation (2026-05-28). Whether weights are posted and their license were not checked; 'open weights' is an inference from the public repo. Benchmark claims not independently verified.
Sources
This record was checked against its sources on 6 October 2026. How we check