I'm particularly excited by Atlas's ability to reconstruct scenes from a very small number of input image -- was able to create this flythrough of London's Natural History Museum by combining 3 input images I found from completely separate sources on Google Images
our new model Atlas also happens to be a capable text-to-image generator, providing it with a strong foundation of “world knowledge”...
this generalizes across both single and multi view domains, meaning you can step into even highly stylized scenes with precise 3D control
These capabilities combine to form an extremely powerful pipeline for sparse 3D reconstruction.
This is a walkthrough of a Gaussian splat of 5 different rooms built from only 30 input images, a quick and easy cellphone camera capture.
Before, accurately reconstructing this entire multi room space would have required many hundreds of images.
Atlas is an autoregressive diffusion model built from the ground up for the task of "next frame prediction".
It is simultaneously a world class method for camera-controlled video generation, novel view synthesis, and sparse 3D reconstruction.