Instead of choosing between training a world model and training a language conditioned robot policy, why not do both? LDA-1B s a new foundation model that is trained on 30,000 hours of human and robot interaction data.
Part of the secret is that LDA-1B jointly learns forward dynamics, action prediction, and visual forecasting, all in a structured DINO latent space which avoids the pitfalls of redundant pixel-level prediction which isn’t necessarily aligned robot action. This approach works on both dexterous hands and simple robot grippers; it also generalizes across objects, tasks, and scenes.
@JiangranLyu joins us to explain. Learn more on Episode 98 of RoboPapers, with
@micoolcho and
@chris_j_paxton!