From 32 input images to real-time flight through
@nvidia's Voyager headquarters.
Trained on NVIDIA Blackwell GPUs, Atlas uses these images as 3D spatial context to generate new views, letting you explore with pixel-perfect camera control.
Take a look around.
Show more
Atlas beta access is rolling out to users, some ideas for API developers
- give Astra "eyes" in Blender, let it place and render new camera views with Atlas
- 3D reconstruction (this video uses 9 image inputs of my studio to make a scene, any Zillow apartment works)
- graybox-to-world, use depth inputs and generate multi-view renders of a scene
Show more
I made a little gaussian splat game demo set in a world created in World Labs Marble, using three.js + spark + crashcat + navcat, and I made a video about it!
On the
@theworldlabs youtube channel now:
+ project github:
My interest in game programming started from a place of curiosity after watching videos like SimonDev's three.js game tutorials. I tried to create something of the same spirit for this one :)
It's a tour of the building blocks and concepts that come together to build a browser game inside a gaussian splat scene. It covers:
- rendering gaussian splats with three.js + spark with level of detail
- splat collision mesh generation
- splat navigation mesh generation
- adding a first person character controller
- baking light probes for a gaussian splat to get environment lighting on characters
- casting shadows from characters on a proxy mesh
- character pathfinding + navigation mesh queries for quest ribbons
- a silly prototype fetch quest in the game world to find the missing spaceship keys
I don't think the game will win any awards 😛 but I hope it is a fun & useful game technology tour. Enjoy!
Show more
atlas lets you walk around any zillow listing in realtime
One of my favorites: Sparse reconstruction powered by Atlas.
The full scene was reconstructed from only 19 photos I took with my phone.
World Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence:
LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century.
The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just three.
In this conversation, they get into the slow motion shot from The Matrix that took hundreds of cameras and now takes three iPhones, the overnight Slack message that made them bet the company in five seconds, why robotics is bottlenecked on data rather than chips, and the case that new view prediction is AI-complete.
00:00 Intro
01:50 The Matrix slow motion scene now takes three iPhones
02:48 Why new view prediction is the primitive
07:10 Unifying generation and reconstruction
11:15 Gaussian splats became the bottleneck
14:17 Dense capture used to mean 300 photos
17:30 Why reconstruction needs generation to fill the gaps
18:44 The LLM lesson image models missed
23:39 The video that made them go all in
28:04 3D design is 95% revisions
30:50 The problem in robotics is data, not chips
32:48 Why a robot policy can't be trained like an image model
34:44 When the simulator becomes the planner
36:45 Frozen time required footage full of movement
40:57 Why new view prediction is AI-complete
42:43 Nature gave animals eyes but not trees
YouTube:
@drfeifei @jcjohnss @BenMildenhall @theworldlabs @martin_casado
Show more
Atlas is our scalable foundation model that can perceive, generate, reason, and interact with virtual and physical worlds.
We're opening up early access in the coming weeks.
Learn more and sign up here:
Show more
All of this is powered by one unified architecture: a multimodal autoregressive diffusion transformer, pretrained from scratch.
This foundation blends the best of modern LLMs and video models, benefiting from the architectural, algorithmic, and systems advances from both areas.
Show more
For robotic simulation, Atlas reconstructs a space from just a few photos and generates the photorealistic RGB and depth data any robot's sensors would observe on any trajectory.
Robots can now be trained and tested in far more spaces. Until now, scanning spaces like these required expensive equipment and time-consuming capture.
Show more
Atlas understands both the spatial structure of the world and how it evolves over time.
This lets us turn a handful of ordinary cameras into a "bullet time" multiview capture studio, capturing dynamic moments frozen in time.
Show more
Atlas also outputs explicit 3D from one to many input images, beating top open source reconstruction models.
Passing more images gives Atlas more context: the more it sees, the less it imagines.
Show more
By positioning multiple input views within the model's spatial context, Atlas allows you to generate image and video frames with precise control.
Here, we hand-designed a camera trajectory to generate a 1 minute video at 1440p resolution from seven reference images.
Show more
We pre-trained Atlas from scratch to take multimodal inputs, including camera movement, and turn it into 3D grounded views.
Atlas puts you in control: direct the views, reconstruct real spaces with your inputs, and build explorable worlds.
Read more:
Show more
Introducing Atlas:
The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.
Model the world, move the camera, and simulate space & time.
Show more
The next generation of robotics needs models that can understand, simulate, learn from, and act across virtual and physical worlds.
We’re hiring a Sr Business Development Lead, Robotics & Physical AI to help bring this technology into real-world applications 🤖🚀
Show more
Viewing these worlds at full scale on an LED stage is mesmerizing.
From 16K LED volumes to real-time production workflows, here are a few ways Marble is being used in virtual production and VFX ↓
Show more
270+ builders joined us in LA for the Worlds in Action Hackathon.
In one weekend, 64 teams built projects ranging from immersive storytelling and mixed reality games to filmmaking tools, tabletop worlds, and audio-reactive experiences.
Check out what they built ↓
Show more
To scale robot intelligence, we must scale the worlds in which robots learn. 🤖🌎
Spatial intelligence is the next frontier for AI.
During its Q2 earnings remarks, Google highlighted growing demand for AI infrastructure from robotics and spatial intelligence companies, including World Labs. ↓
Show more