AI Research Atlas

Qwen-VLA (vision-language-action generalist)

Alibaba (Qwen) · 28 May 2026

Unified vision-language-action model on Qwen3.5-4B plus a 1.15B flow-matching action decoder, covering manipulation, navigation and trajectory prediction.

One generalist trained jointly on robot manipulation, human egocentric, simulation and navigation data. Qwen says it matches or beats specialists fine-tuned per benchmark. The June Qwen-Robot Suite ships separate navigation and manipulation models; their link to Qwen-VLA was not checked.

Date
Thursday, 28 May 2026
Lab
Alibaba (Qwen)
Kind
model
Access
open weights

Figures

MeasureValueMeasured by
Technical reportarXiv 2605.30280
Revised v2 on 2026-06-01
company

Date is GitHub repo creation (2026-05-28). Whether weights are posted and their license were not checked; 'open weights' is an inference from the public repo. Benchmark claims not independently verified.

Sources

  1. github.com/QwenLM/Qwen-VLA
  2. arxiv.org/abs/2605.30280

This record was checked against its sources on 6 October 2026. How we check

Related