From a single casual smartphone video, build a free-viewpoint "moving 3D human" — no studio, no camera rig required.
Title: 4DAnyone: Create Anyone in 4D from a Casual Monocular Video
URL:
A framework that turns an uncalibrated monocular video into multi-view-consistent video, then into 4D Gaussian Splatting. Three highlights.
📦 Highlight 1: Compress reference context from O(N) to O(1) (RCP)
4D reconstruction needs dozens of views, but conditioning on generated views bloats the context and weakens guidance. Reference Context Packing packs references into fixed-length mixed-resolution context, keeping compute constant as views grow.
🔄 Highlight 2: Rotate view groups by noise level (TCR)
Independently denoised view groups can't share info, so structure drifts. Target Context Routing cyclically shifts groups at high noise to propagate global structure, then fixes groups at low noise to stabilize detail. The switch at t/T=0.2 was optimal.
🎮 Highlight 3: Game-engine data for in-the-wild generalization
38k in-house multi-view videos (318 actors, 24 cameras) plus real light-stage and monocular data, trained via a 3-stage curriculum. It hits PSNR 24.33 on generation consistency, clearly beating prior methods.
Photorealistic 4D avatars from phone video in ~40 min — free-viewpoint capture for everyone.
#
4DReconstruction# #
GaussianSplatting#