Ah!! We've been trying to get this working for months.
Basically, we can now capture a person and the objects they're interacting with in 3D from regular video. Motion for both, plus approximate object shape.
I've been excited about this for a while because so much of what people do involves stuff. If you want a robot to learn how to use a tool, or to understand someone's tennis swing, you probably want to know what the tool or racket is doing too.
Same goes for props in animation and everyday tasks in occupational therapy. There's a lot here.
Still imperfect, but this week it started working well enough that I wanted to show you.