Internet video is now robot training data! ๐
@perceptroninc, the startup from the ex-Meta team behind the Chameleon multimodal models, just released Isaac 0.5, an open-source embodied foundation model.
The same weights answer questions about a video, point to objects, track them over time, report task progress, and generate the robot's actions. What it learns from perception directly shapes what it does.
To make a big model run fast enough for control, they built on something called Null Experts: each token can pick how many experts it activates, so the model spends more compute on hard parts of a scene and coasts through the easy ones.
โ 36B-parameter sparse model, trained across 35+ robot systems and 3T multimodal tokens
โ 62.6 on ScreenSpot-Pro grounding vs 54.8 for the strongest Qwen3-VL run
โ Does it at ~8.5ร lower inference cost, 26.9 TFLOP per request against 228.4
โ On a chess-manipulation task, one epoch on a single episode cut action loss 10.5ร
Perception and action sharing one brain, with compute that flexes to the task, that's a genuinely different shape for a robot model.
And it's fully open, weights, training code, inference via LeRobot.
Here's the blog:
~~
โป๏ธ Join the weekly robotics newsletter, and never miss any news โ