Robot foundation models look great in simulation, but fail the moment the camera angle or lighting changes. Turns out they might just be cheating on looks.
Title: Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models
URL:
📝 Overview
This paper proposes Latent Interface Training (LIT), a method that breaks the "vision-action shortcut" where models exploit task-irrelevant visual cues that just happened to correlate with actions in the training distribution.
❗ Problem it solves
Robot foundation models perform well in-distribution but degrade sharply under visual distribution shift — different camera viewpoints, lighting, or sensor noise.
⚙️ Methodology
It's a two-stage recipe: first pretrain the action expert using only pose goals, no images at all, then route all visual information through 100 latent tokens supervised with a pose-reconstruction loss.
🤖 Use cases
It applies to 4 different robot foundation models — π0.5, MolmoAct2, FAST-WAM, and ImageWAM — making it a genuinely framework-agnostic technique.
📊 Results
On LIBERO-Plus out-of-distribution evaluation, every model improved by 3.87 to 10.70 points. In real-world robot manipulation, it gained +16.7 points under lighting changes, and up to +60 points on a specific task.
Relying on real spatial information instead of surface appearance feels like a direct path to more trustworthy robots in the field.
#
Robotics# #
VLA#