3D has always felt essential for visual intelligence that truly works in the real world.
The challenge is making it useful in a simple and scalable way.
Cambrian-P shows a surprisingly strong signal: adding pose grounding to video understanding substantially improves spatial reasoning, recognition, and even general video QA.