註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture 投稿は個人の意見です。
加入 May 2026
280 正在關注    413 粉絲
Robot foundation models look great in simulation, but fail the moment the camera angle or lighting changes. Turns out they might just be cheating on looks. Title: Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models URL: 📝 Overview This paper proposes Latent Interface Training (LIT), a method that breaks the "vision-action shortcut" where models exploit task-irrelevant visual cues that just happened to correlate with actions in the training distribution. ❗ Problem it solves Robot foundation models perform well in-distribution but degrade sharply under visual distribution shift — different camera viewpoints, lighting, or sensor noise. ⚙️ Methodology It's a two-stage recipe: first pretrain the action expert using only pose goals, no images at all, then route all visual information through 100 latent tokens supervised with a pose-reconstruction loss. 🤖 Use cases It applies to 4 different robot foundation models — π0.5, MolmoAct2, FAST-WAM, and ImageWAM — making it a genuinely framework-agnostic technique. 📊 Results On LIBERO-Plus out-of-distribution evaluation, every model improved by 3.87 to 10.70 points. In real-world robot manipulation, it gained +16.7 points under lighting changes, and up to +60 points on a specific task. Relying on real spatial information instead of surface appearance feels like a direct path to more trustworthy robots in the field. #Robotics# #VLA#
顯示更多