登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture 投稿は個人の意見です。
参加 May 2026
280 フォロー中    415 ファン
Robot foundation models look great in simulation, but fail the moment the camera angle or lighting changes. Turns out they might just be cheating on looks. Title: Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models URL: 📝 Overview This paper proposes Latent Interface Training (LIT), a method that breaks the "vision-action shortcut" where models exploit task-irrelevant visual cues that just happened to correlate with actions in the training distribution. ❗ Problem it solves Robot foundation models perform well in-distribution but degrade sharply under visual distribution shift — different camera viewpoints, lighting, or sensor noise. ⚙️ Methodology It's a two-stage recipe: first pretrain the action expert using only pose goals, no images at all, then route all visual information through 100 latent tokens supervised with a pose-reconstruction loss. 🤖 Use cases It applies to 4 different robot foundation models — π0.5, MolmoAct2, FAST-WAM, and ImageWAM — making it a genuinely framework-agnostic technique. 📊 Results On LIBERO-Plus out-of-distribution evaluation, every model improved by 3.87 to 10.70 points. In real-world robot manipulation, it gained +16.7 points under lighting changes, and up to +60 points on a specific task. Relying on real spatial information instead of surface appearance feels like a direct path to more trustworthy robots in the field. #Robotics# #VLA#
もっと見る