Qwen just stepped into autonomous driving! đ
Qwen-Drive-1.0-4B is a vision-language foundation model that handles 3D perception, driving VQA, and motion planning in one framework, with the Qwen3.5-4B backbone left fully unmodified. Apache 2.0. đ¤
âī¸ Two plug-in modules do the driving: a BEV head for 3D perception, and a flow matching Planning Expert for trajectories.
đ Leads driving VQA across the board: 77.8 on LingoQA, lowest Ego3D distance error, and 41.3 on causal reasoning where others score under 5.
đ The RL planner hits 90.7 PDMS on NAVSIM, ahead of AutoVLA and SpanVLA.
đ§ No catastrophic forgetting: general benchmarks stay on par with the base model.