“World in World: Explore the World with World Models”
Video world models can generate long rollouts, but controlling them from new viewpoints usually needs task-specific training or adapters.
This paper instead turns source frames, geometry, and past generated states into visual evidence that a frozen world model can directly read through self-attention.
This then gives training-free camera-controlled rerendering with better long-horizon consistency, unseen-view completion, and lower camera error than prior methods.