AI Research Atlas

Qwen-Drive-1.0 (4B)

Alibaba (Qwen) · 2 September 2026

Vision-language foundation model for autonomous driving, built on Qwen3.5-4B, unifying 3D perception with visual question answering.

Keeps the pretrained vision-language model's architecture and adds an external bird's-eye-view perception head (3D detection, occupancy, map segmentation) plus a Planning Expert that generates future ego trajectories; staged training mixes driving and general data. Qwen calls it an initial step.

Date
Wednesday, 2 September 2026
Lab
Alibaba (Qwen)
Kind
open-weights
Access
open weights

Blog dated 2026-09-02 per Releasebot; HF Qwen-Drive-1.0-4B repo created 2026-08-27 and GitHub repo 2026-08-25; paper arXiv 2609.00111 submitted 2026-08-31 (16 authors). Open-loop, pseudo-closed-loop and closed-loop results are in the paper; none were reproduced here.

Sources

  1. releasebot.io/updates/qwen
  2. huggingface.co/api/models?author=Qwen&sort=createdAt&direction=-1&limit=60
  3. arxiv.org/abs/2609.00111

This record was checked against its sources on 6 October 2026. How we check

Related