註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

alphaXiv
@askalphaxiv
High fidelity research
加入 November 2023
101 正在關注    56.9K 粉絲
“World in World: Explore the World with World Models” Video world models can generate long rollouts, but controlling them from new viewpoints usually needs task-specific training or adapters. This paper instead turns source frames, geometry, and past generated states into visual evidence that a frozen world model can directly read through self-attention. This then gives training-free camera-controlled rerendering with better long-horizon consistency, unseen-view completion, and lower camera error than prior methods.
顯示更多
0
11
276
41
轉發到社區