A video world model for robot manipulation that actually verifies whether the generated video faithfully follows the prescribed actions has arrived.
DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation
Built on Wan2.2-TI2V-5B, this model predicts future frames with high fidelity from an initial observation, language instruction, and bimanual action trajectory. It ranked 1st among 31 teams on WorldArena 2.0 Track 1 (EWMScore-P: 60.65).
🔷 Highlight 1: PRoPE — Injecting SE(3) Directly into Attention
Instead of compressing actions into generic tokens, the model injects end-effector positions, rotation matrices, and gripper states as SE(3) transformations directly into the attention mechanism. Attention heads are partitioned per arm, and token-wise transforms are applied via Kronecker products to eliminate dependence on global coordinate frames. This yields a remarkable controllability score of 98.55.
🔶 Highlight 2: Depth Branch + Object-Centric Supervision Beyond Visual Plausibility
A two-pronged approach tackles the fundamental problem that RGB loss alone cannot constrain surface ordering or object extent. A lightweight depth branch (the final M blocks replicated with one-way cross-attention) enforces geometric consistency, while SAM3 masks combined with Gram matrix constraints from a frozen V-JEPA teacher preserve temporal coherence of manipulated objects — ensuring contact-local errors remain influential despite large static backgrounds.
🟣 Highlight 3: DMD Distillation Compresses Multi-Step into Few-Step
Distribution-Matching Distillation (DMD) combining KL divergence and a non-saturating GAN loss drastically reduces inference steps. Trained on over 6,000 hours of diverse data spanning Ego4D, AgiBot World 2026, and RoboTwin 2.0, the model balances broad visual priors with precise action grounding.
Robot world models have taken a decisive step from "looks realistic" to "moves correctly."
#
Robotics# #
WorldModel#