“In-Context Robot Learning with VLM Agents”
Most robots need lots of task-specific training before they can do something new. This paper shows you can instead just show a general VLM what to do and let it adapt directly from context.
GPT-Policy essentially turns robot learning into prompting. So they gave the frozen VLM an example of the task, even just a human video with no robot action labels, and it can translate what it sees into robot actions on the fly.
This brings few-shot learning into the physical world, where teaching a robot a new behavior could look more like showing it an example than retraining a policy.