Qwen-Drive-1.0 (4B)
Vision-language foundation model for autonomous driving, built on Qwen3.5-4B, unifying 3D perception with visual question answering.
Keeps the pretrained vision-language model's architecture and adds an external bird's-eye-view perception head (3D detection, occupancy, map segmentation) plus a Planning Expert that generates future ego trajectories; staged training mixes driving and general data. Qwen calls it an initial step.
- Date
- Wednesday, 2 September 2026
- Lab
- Alibaba (Qwen)
- Kind
- open-weights
- Access
- open weights
Blog dated 2026-09-02 per Releasebot; HF Qwen-Drive-1.0-4B repo created 2026-08-27 and GitHub repo 2026-08-25; paper arXiv 2609.00111 submitted 2026-08-31 (16 authors). Open-loop, pseudo-closed-loop and closed-loop results are in the paper; none were reproduced here.
Sources
- releasebot.io/updates/qwen
- huggingface.co/api/models?author=Qwen&sort=createdAt&direction=-1&limit=60
- arxiv.org/abs/2609.00111
This record was checked against its sources on 6 October 2026. How we check