AI Robotics research is showing that you can build increasingly powerful robot models by pretraining on the large preexisting corpus of web video data without collecting vast amounts of new data
Visual dynamics understanding increases with more compute and that directly translates into more performant robot foundation models
Does scaling pre-training on general web video improve a complex manipulation task in real deployment?
We scale model size and pre-training compute, and test on one industrial task.
Yes. The better a pre-trained model predicts web video, the better its post-trained policy. 🧵