This is what robotics policy training is heading to.
How it works now: engineer - reward function - mass training - success
But Nvidia ASPIRE is more like: Receive a new task - Execute the policy - Evaluate the outcome - Identify failure patterns - Optimize the policy - Repeat - Master the task
Self-improvement for robotics model.