Register and share your invite link to earn from video plays and referrals.

Jihan Yang
@jihanyang13
@amilabs; Prev. @NYU_Courant @HKUniversity; Researcher in Deep Learning, Computer Vision.
522 Following    1.3K Followers
see you at ICML in Seoul and the AMI × SBVA mixer on thursday evening! looking forward to meeting everyone :)
I recently joined and we're looking for software engineers to build a AI infra team on-site here in Singapore If you're in Singapore tired of big-tech satellite offices or boring CRUD apps, and want to work on cutting edge AI infra, please reach out!
Show more
Modern text-to-image models are increasingly powered by large pretrained LLMs. But there is a curious mismatch: the LLM typically encodes the prompt only once, while the evolving noisy latent states are handled entirely by a newly trained generative backbone. Can pretrained multimodal prior participate in the denoising process? Introducing RepFusion. (1/12) 📄 🌐
Show more
introducing PaintBench 🎨 a deterministic eval protocol + benchmark for precise visual editing unlike most generative model evals, no judge model in the loop see @itskaixu's 🧵⬇️ for technical details. I'll provide more context / commentary about why it's exciting here
Show more
Can MLLMs actually track what's happening in a video? Introducing VSTAT 🎯, our new benchmark for visual state tracking. The tasks are simple: count cups, read typed words, count page flips. Humans solve them easily. MLLMs don't. 🧵 [1/11]
Show more
Image editing models can put you on the Moon, but can they precisely move a circle right by 50 pixels? 📐 Introducing 🎨PaintBench: a foundational eval of visual editing operations with only one right answer. The highest-performing model (@NanoBanana 2) reaches only 17.1%.
Show more
Pose is so essential for grounding humans in the 3D world -- and the same applies for MLLMs. ​Happy to share Cambrian-P, a year-long collaboration with NYU/FAIR. We introduced a simple pose token to MLLMs, and it just works!
Show more
We need a proxy task that goes beyond simple VQA to explore the model's potential for understanding the real world (especially spatial intelligence), Cambrian-P is an excellent attempt!
3D has always felt essential for visual intelligence that truly works in the real world. The challenge is making it useful in a simple and scalable way. Cambrian-P shows a surprisingly strong signal: adding pose grounding to video understanding substantially improves spatial reasoning, recognition, and even general video QA.
Show more
What is the most elegant way to give MLLMs spatial awareness? Instead of adding heavy 3D modules, we let the model learn a simple question: “Where am I, and where am I looking?” Introducing Cambrian-P, a new learning paradigm for video understanding. (1/n)
Show more
Check out Cambrian-P, a must-read if you work on video/3D MLLMs or spatial understanding. Key takeaway: camera pose, which naturally describe the relationships between isolated frames, turns out to be highly effective for enhancing video understanding and spatial intelligence!
Show more
Camera pose matters for video understanding! Today's MLLMs excel at recognizing activities, but still struggle with the underlying space and ego/object dynamics in video. We trace this gap to a missing piece: camera pose. Introducing Cambrian-P: a multimodal LLM natively grounded in camera pose. (1/n)
Show more