Register and share your invite link to earn from video plays and referrals.

Saining Xie
@sainingxie
cofounder & chief science officer at @amilabs | faculty @nyu_courant | prev: @googledeepmind @meta (fair) @ucsandiego | ynwa
1.6K Following    40.2K Followers
I just graduated from Stanford, and surprisingly, I feel happier, more energized, and more excited than ever. The SF bubble can be an echo chamber. In search of more original, creative, and grounded thinking, I’m moving to New York to join @amilabs alongside @ylecun, @sainingxie, and an incredible team researching on world modeling and physical intelligence. Excited to branch out and meet people from all walks of life in NYC. Please reach out :)
Show more
I am excited to share that I am joining @amilabs as Director of Research, Paris, working with @ylecun and an exceptional founding team. Further progress in machine intelligence will require not only scaling foundation models and large-scale engineering, but also new ideas and breakthroughs. This is what makes AmiLabs such a unique and exciting place to build. Its focus on world modeling — systems that can learn richer representations of the real world, reason, plan, and learn from interaction — is a long-term research direction I am deeply excited about. I am particularly looking forward to working with extraordinary co-founders @sainingxie, @lxbrun, @michaelrabbat, @pascalefung, @mavenlin, @laurentsolly, and the amazingly talented AmiLabs team, to helping build the research organization and to the journey ahead.
Show more
After almost 7 incredible years at Meta, I’m excited to share that I’ll be joining the team at AMI Labs in New York to help build real-world models! AMI’s mission is something that has been close to my heart since my early days at FAIR. Much of my work over the years has focused on building perception and generation models that leverage the temporal structure of videos to predict future visual states, much like how language models leverage the structure of language to predict the next token. Along the way, this work helped advance a range of visual tasks, including physical reasoning, action anticipation, image animation, and video generation, while also contributing to some of the largest video datasets and benchmarks in the world. Now that language models have demonstrated remarkable capabilities in coding, math, and reasoning, I believe it’s the right time to push toward the next frontier: models that can achieve similar impact across a much broader range of real-world tasks. Leaving Meta has been one of the hardest decisions of my life. It was my first job after grad school, and in many ways, I grew up there. Many of the people I met there became close friends, mentors, and collaborators. When I joined, billion-parameter models were just beginning to emerge, vision systems were largely built separately for each task, and video generation models were hard to control and could produce little more than low-resolution clips. Needless to say, a lot changed over the years, and I was incredibly fortunate to have a front-row seat to that progress and to contribute along the way. I had the privilege of working on projects spanning multimodal modeling, representation learning, and video generation, including Ego4D, Omnivore, ImageBind, Emu Video, Llama-3, and MovieGen. None of this would have been possible without the exceptionally talented and kind collaborators I had the opportunity to work alongside. I’ll always be thankful to Meta for the opportunities, growth, learning, and most importantly, the friendships that will stay with me for a lifetime. Looking ahead, I’m deeply grateful to @sainingxie @michaelrabbat Min Lin @jingli9111 @ylecun @lxbrun and the entire AMI team for the opportunity to join this ambitious mission. I’m excited to help build the next generation of real-world models and push toward AI systems that better understand and interact with the physical world. Onwards! 🚀
Show more
A triple life update: 1. I recently moved from NYC to London 2. I’m starting a new lab at @imperialcollege 3. I’ve joined my friends and mentors at @amilabs as a founding Member of Technical Staff It's time to build 🏗️🧑‍🍳
Show more
Modern text-to-image models are increasingly powered by large pretrained LLMs. But there is a curious mismatch: the LLM typically encodes the prompt only once, while the evolving noisy latent states are handled entirely by a newly trained generative backbone. Can pretrained multimodal prior participate in the denoising process? Introducing RepFusion. (1/12) 📄 🌐
Show more
VLA-JEPA just dropped in LeRobot 🤖 What makes this model special is that it does not just learn what action to take from a given observation, it also leverages a JEPA world model to learn action-relevant dynamics. During training, the VLA leverages V-JEPA2 by conditioning its predictor. This clever trick adds a world modeling objective to the training, which also allows pretraining on human videos. At inference, the world model is dropped entirely, keeping only a standard VLA architecture: Qwen backbone and action head. The demo here was only fine-tuned on 13 examples, showing great pretraining capability and running in real time on @NVIDIARobotics DGX Spark! VLA-JEPA is the first world model to be ported to LeRobot, and I feel like it won't be the last 🚀 @Thom_Wolf @ClementDelangue
Show more
0
32
1.4K
191
Forward to community
Congrats @thegautamkamath!!
Honoured that our 2016 paper, Robust Estimators in High Dimensions without the Computational Intractability, w/ Ilias Diakonikolas, Daniel Kane, Jerry Li, Ankur Moitra, Alistair Stewart, was awarded the 2026 Gödel Prize This is the highest award for papers in theoretical CS. 1/7
Show more
VSTAT highlights the substantial perceptual gap between humans and MLLMs, but it goes far beyond that. Its diverse tasks are designed not merely to assess simple pixel-space tracking, but to evaluate how well models capture and understand evolving world states in the latent space of videos. Text is only one way to probe this capability, and we are excited to see future evaluations explore new modalities such as pixels, actions, and beyond! Working on this benchmark has been a lot of fun along the way—huge shout-out to my amazing collaborators!
Show more
how does the brain build and track an internal state of the world from (possibly incomplete and noisy) visual observations? i believe visual state tracking will be the grand challenge for vision in the coming years, and i hope this benchmark can be a useful starting line. enjoy!
Show more
Can MLLMs actually track what's happening in a video? Introducing VSTAT 🎯, our new benchmark for visual state tracking. The tasks are simple: count cups, read typed words, count page flips. Humans solve them easily. MLLMs don't. 🧵 [1/11]
Show more
State tracking is a core pillar of video understanding: it requires identifying entities and events, and mapping how their states evolve over time. Frontier multimodal models are surprisingly bad at it, so we built a benchmark to measure it. Meet VSTAT!
Show more
I’m thrilled to share that I have joined AMI Labs as a Member of Technical Staff! 🚀 We are building the future of AI in the physical world and actively hiring from all our offices. Reach out to me if you are interested in research roles in our Paris office!
Show more
Image editing models can put you on the Moon, but can they precisely move a circle right by 50 pixels? 📐 Introducing 🎨PaintBench: a foundational eval of visual editing operations with only one right answer. The highest-performing model (@NanoBanana 2) reaches only 17.1%.
Show more
@_amirbar thanks! in the second paper ( we used your (and RAE's) recipe and it worked.
What is the most elegant way to give MLLMs spatial awareness? Instead of adding heavy 3D modules, we let the model learn a simple question: “Where am I, and where am I looking?” Introducing Cambrian-P, a new learning paradigm for video understanding. (1/n)
Show more
📸latest in our cambrian series: cambrian-p, p for pose. i think pose is probably the minimal sufficient 3d signal (and it’s easy to get!) that we need for robust video multimodal models -- jointly modeling frames and pose turns image sequences into a globally grounded structure.
Show more
Camera pose matters for video understanding! Today's MLLMs excel at recognizing activities, but still struggle with the underlying space and ego/object dynamics in video. We trace this gap to a missing piece: camera pose. Introducing Cambrian-P: a multimodal LLM natively grounded in camera pose. (1/n)
Show more
check out RAEv2 led by Jas. through extensive exps, we found some really intriguing behaviors showing why strong representation encoders are key for pixel decoders. spoiler: it’s not about hillclimbing fid; new metrics like ep@fid-k/fdr^k show there’s a lot more left to explore!
Show more
In Oct last year, Representation Autoencoders provided an elegant solution to unified tokenization for understanding and generation. Today we make them a bit more simple. a bit more general. Result: >10x faster convergence, better reconstruction, better generation. And yes we test them on T2I and world models :) Introducing RAEv2
Show more
So, today is my last day at Meta... After I finished my PHD, I moved from Oxford to New York to join FAIR and work on generative models for video. At the time, the SOTA could do little more than generate blurry GIFs. In the subsequent three years, our team went to complete my proudest research achievements so far, pushing the state of the art a few times with Emu Video and Movie Gen. For anyone interested, I summarized the amazing progression in the field and our work in a lecture at Stanford last year I worked with an amazing close team throughout in @imisra_, @_rohitgirdhar_, @mannat_singh, @quduval, Xi Yin, @smnh_azadi , @rssaketh , @arunmallya , @deviparikh , Ce Liu, all of whom are close friends now. I’m gonna be cheering the team on from outside. After pushing the frontier of video generation for a few years, its time for a change. I’m excited to announce that I will be joining the super talented team at @amilabs as a member of the technical staff with @sainingxie I’m so excited to join this talented team, to push the frontier of world models, and to do exciting new research. This is gonna be a fun ride.
Show more
on the topic of favorite filters… first emergence, mostly accidental
I remember the days when a successful experiment was one that produced cool filters