A humanoid can nail the grasp and still fail the task.
Opening a low drawer isn't just the hand. It's the stance, the weight shift, and staying balanced through the whole reach. Hand trajectories alone don't capture any of that.
That's what makes
@GenrobotAI's whole-body mesh approach interesting. Using vision alone, they reconstruct full-body motion from first-person video, reporting ~3 cm whole-body error (about 2 cm upper body, 3.5 cm lower).
The hard part: an ego camera barely sees the body. Hands vanish behind objects, arms cross, fast motion blurs, and long tasks drift. Nailing one clean frame is easy. Keeping a whole task consistent over time is the real problem.
And they go past the mesh itself. The pipeline aligns hand-object contact (when it starts, holds, and releases) and runs automated quality checks to output model-ready data. A robot doesn't just need a pose. It needs how motion, contact, and object interaction unfold together.
Their reported capacity: 100K hours of this a month.
The story is the combination: detailed reconstruction plus a scalable pipeline to produce usable data. As humanoids move past the tabletop, whole-body data could become a core part of the next training-data paradigm.