๐ค What changes when a robot doesn't just act, but thinks ahead and predicts the future before it moves? Here's an embodied foundation model trained on over 37,000 hours of data.
Title: GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture
URL:
GigaBrain-0.7 proposes a three-system architecture that cleanly separates yet integrates understanding, prediction, and action. Three highlights stand out.
๐ง System 2: Understanding and planning
Built on PaliGemma2 (3B), it interprets the current visual scene and long-horizon instructions, using chain-of-thought reasoning to break "hang it on a hanger" into a subtask like "first pick up the red hat."
๐ฎ System 3: Prediction and evaluation
A GigaWorld-1-based model (5B) predicts future visual states as video, extracting the final frame as a "subgoal image." It also scores task progress as a value signal that guides the action system.
๐ฆพ System 1: Action and control
A Mixture-of-Transformers setup generates continuous action chunks, using a unified embodiment representation across 16 robot types for cross-robot generalization.
What stands out most is that larger-scale pretraining consistently lowered validation loss, with signs of genuinely emergent embodied capabilities showing up along the way.
#
Robotics# #
EmbodiedAI#