Register and share your invite link to earn from video plays and referrals.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
Joined May 2026
280 Following    415 Followers
Robot foundation models look great in simulation, but fail the moment the camera angle or lighting changes. Turns out they might just be cheating on looks. Title: Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models URL: ๐Ÿ“ Overview This paper proposes Latent Interface Training (LIT), a method that breaks the "vision-action shortcut" where models exploit task-irrelevant visual cues that just happened to correlate with actions in the training distribution. โ— Problem it solves Robot foundation models perform well in-distribution but degrade sharply under visual distribution shift โ€” different camera viewpoints, lighting, or sensor noise. โš™๏ธ Methodology It's a two-stage recipe: first pretrain the action expert using only pose goals, no images at all, then route all visual information through 100 latent tokens supervised with a pose-reconstruction loss. ๐Ÿค– Use cases It applies to 4 different robot foundation models โ€” ฯ€0.5, MolmoAct2, FAST-WAM, and ImageWAM โ€” making it a genuinely framework-agnostic technique. ๐Ÿ“Š Results On LIBERO-Plus out-of-distribution evaluation, every model improved by 3.87 to 10.70 points. In real-world robot manipulation, it gained +16.7 points under lighting changes, and up to +60 points on a specific task. Relying on real spatial information instead of surface appearance feels like a direct path to more trustworthy robots in the field. #Robotics# #VLA#
Show more