Qwen just stepped into autonomous driving! ๐
Qwen-Drive-1.0-4B is a vision-language foundation model that handles 3D perception, driving VQA, and motion planning in one framework, with the Qwen3.5-4B backbone left fully unmodified. Apache 2.0. ๐ค
โ๏ธ Two plug-in modules do the driving: a BEV head for 3D perception, and a flow matching Planning Expert for trajectories.
๐ Leads driving VQA across the board: 77.8 on LingoQA, lowest Ego3D distance error, and 41.3 on causal reasoning where others score under 5.
๐ The RL planner hits 90.7 PDMS on NAVSIM, ahead of AutoVLA and SpanVLA.
๐ง No catastrophic forgetting: general benchmarks stay on par with the base model.