What is the most elegant way to give MLLMs spatial awareness?
Instead of adding heavy 3D modules, we let the model learn a simple question:
“Where am I, and where am I looking?”
Introducing Cambrian-P, a new learning paradigm for video understanding. (1/n)