How do we make robot policies robust to rare but high-impact failures?
Video #
World# #
Models# (WMs) are rapidly becoming a powerful tool for robotics, enabling policy evaluation and improvement by "imagining" future outcomes. But there's a catch: these imagined futures are typically nominal samples, making it easy to overlook the rare yet safety-critical events that matter most.
In our new paper, StressDream: Steering Video World Models for Robust Policy Evaluation and Improvement, we explore a simple but powerful idea:
đĄ Instead of passively sampling futures, actively steer world model imaginations toward high-impact yet still plausible scenarios.
StressDream optimizes the initial diffusion noise at inference time, allowing us to generate targeted stress-test scenarios without retraining the world model. This enables:
- More robust policy evaluation by exposing failure modes that random sampling often misses.
- Improved policy optimization by training against challenging but realistic imagined futures.
As generative world models become a foundation for #
Physical# #
AI#, the ability to systematically probe their "long tail" of plausible futures will be increasingly important for building reliable and trustworthy autonomous systems.
đ đ¯đđđđžđŧđ đ¯đēđđž:
đ đ¯đēđđžđ:
Work led by Junwon Seo, with a great set of collaborators: Sushant Veer, Thomas Ran Tian, Wenhao Ding, Apoorva Sharma, Karen Leung, Edward Schmerling, Andrea Bajcsy.
@NVIDIADRIVE @NVIDIAAI
#
Robotics# #
WorldModels# #
PhysicalAISafety# #
AISafety# #
AutonomousSystems# #
RobotLearnin#