Qwen just stepped into autonomous driving! 🚗
Qwen-Drive-1.0-4B is a vision-language foundation model that handles 3D perception, driving VQA, and motion planning in one framework, with the Qwen3.5-4B backbone left fully unmodified. Apache 2.0. 🤖
⚙️ Two plug-in modules do the driving: a BEV head for 3D perception, and a flow matching Planning Expert for trajectories.
📊 Leads driving VQA across the board: 77.8 on LingoQA, lowest Ego3D distance error, and 41.3 on causal reasoning where others score under 5.
🏁 The RL planner hits 90.7 PDMS on NAVSIM, ahead of AutoVLA and SpanVLA.
🧠 No catastrophic forgetting: general benchmarks stay on par with the base model.