Excited to share our latest research introducing Qwen-VLA—a unified Vision-Language-Action model for general embodied intelligence 🤖
By combining Qwen3.5-4B with a 1.15B DiT decoder, it unifies manipulation, navigation, and trajectory prediction into a single framework.
With embodiment-aware prompts, the same Qwen-VLA model can operate across 11 robot embodiments under a unified architecture—covering single-arm, dual-arm, and humanoid platforms without task-specific policy heads—without task-specific architecture forks or separate policy heads.