Register and share your invite link to earn from video plays and referrals.

Sanja Fidler
@FidlerSanja
Associate Professor @UofT, Vice President of AI Research @nvidia, founding member of @VectorInst. Computer vision, deep learning, 3D. Opinions are my own.
496 Following    18.8K Followers
Today, I’m incredibly excited to introduce Veeda AI, which I co-founded with @FidlerSanja and @ZGojcic . Throughout my career, I’ve been fascinated by one question: how can visual generative models impact the physical world? My journey has taken me from images to video to world foundation models. Now, we’re building toward the next frontier: turning these generated worlds into virtual interactive playgrounds where physical intelligence can perceive, reason, act, and ultimately learn within them at scale. That’s our ambition at Veeda and I’m incredibly excited to build this alongside Sanja, Zan, and our amazing team. We’re hiring. Come build it with us.
Show more
I’m incredibly excited to introduce @VeedaAI , which I co-founded with my incredible longtime collaborators @ZGojcic and @HuanLing6 . We started Veeda because we believe robotics will reshape the world—changing how we move people and goods, how we manufacture and build, and how we operate in the physical world. We also firmly believe that the scaling moment for Physical AI will come from robots learning through interaction with the world. Just like humans learn from interaction with their environments by trying, failing and trying again, until we succeed. And just like the capabilities of LLMs were truly unlocked when they started training in interactive environments and not only on vast amounts of internet text. For Physical AI, the central challenge is making this kind of interactive learning possible at scale. The real world is simply not a practical training ground for robots to learn through trial and error. It is unsafe, expensive, and time-consuming. To scale interactive learning, robots will need to learn in simulated reality. At Veeda, our sole mission is to build simulated reality for Physical AI. Our conviction is that this will become the critical infrastructure layer for all areas of robotics. As envisioned in pop culture, this means building “the Matrix” for Physical AI, where, through a virtual embodiment, robot intelligence can interact with an entirely virtual world in a million possible ways, learning from its own mistakes. We believe that World Models are the foundational technology that can make this possible. These generative models learn from massive amounts of sensor and physical-world data to reach the quality, diversity, and physical realism of the simulated reality that robots require. I could not be more excited to take on this challenge and build this transformative technology alongside an extraordinary team of engineers and researchers who have spent years working on some of the hardest problems across simulation, generative AI, 3D, robotics, and embodied intelligence. Together with our investors at @khoslaventures and @radicalvcfund , we intend to push the frontier of this technology and build the foundational infrastructure for interactive robotics learning and evaluation at scale. Come build this future with us. And to roboticists, help us understand what you need, we’re building this for you.
Show more
0
95
1.1K
84
Forward to community
Check out #InstantNuRec#, which performs feed-forward 3D Gaussian reconstruction for driving scene simulation within seconds. We have now made this into the official NVIDIA #Omniverse# NuRec product. Code and models are available.
Show more
Actually, both of these two NVIDIA live demos at CVPR are powered by flashdreams!
Following recent World-Action Model results in robotics, the same ~2B OmniDreams single-view backbone can be fine-tuned into a driving policy. In preliminary closed-loop results, it reduces collision from 6.9% to 4.2% when compared with Alpamayo 1.5, while having roughly 5x fewer parameters.
Show more
Alpasim now lets you test your AV policies with an interactive world model!
It’s been a while since I posted here, but I’m very excited to share what our team at @nvidia has been building over the past year! After a year of active development, we’re getting ready to release SIL-Wheel to the world: a one-stop shop platform for data-centric workflows in large-scale video model training. Built by researchers, for researchers, SIL-Wheel brings together search, curation, annotation, evaluation, and analysis for large video datasets in one centralized framework. Want a sneak peek before the official release? Come by the NeXD26 Workshop @CVPR tomorrow at 10:30!🚀
Show more
Real time world model NVIDIA OmniDreams now open sourced! If you are at CVPR, we invite you to also check out a live demo you can try out at the NVIDIA booth.
🚀 What if physical AI policies could interact with generated worlds in real time? Introducing OmniDreams, a generative world model for closed-loop autonomous vehicle simulation. Tech report, code, models, and data samples are available now. Project: Code: Model: Join the #omnidreams# discord channel:
Show more
🚀 What if physical AI policies could interact with generated worlds in real time? Introducing OmniDreams, a generative world model for closed-loop autonomous vehicle simulation. Tech report, code, models, and data samples are available now. Project: Code: Model: Join the #omnidreams# discord channel:
Show more
Super excited to open source Flashdreams, acceleration library for world models! No restrictions on usage, Apache 2 license. World models are the next frontier for video generation and paramount to Physical AI. Flashdreams yields out of the box speed up across world models
Show more
World models are moving beyond offline generation towards interactive, real-time experiences. Introducing ⚡FlashDreams⚡: an open-source high-performance inference and serving library built for autoregressive world models: 🔥 Up to 3.10× faster LingBot-World inference 🔥 Up to 2.12× faster Self-Forcing inference 🔥 Up to 1.40× faster Wan2.1 inference 🔥 8 integrated models 🔥 Multi-GPU, streaming, low-latency serving 🔥 Agentic skills that teach you how to use it FlashDreams is designed for a new generation of AI systems that continuously evolve over time while responding to user interactions. It powers applications across robotics, autonomous vehicle simulation, gaming, and virtual worlds. Github: Docs: Research page: Join the #flashdreams# Discord channel at FlashDreams is also the runtime backbone behind NVIDIA OmniDreams ( 1/n #AI# #WorldModels# #FastInference# #PhysicalAI# #OpenSource# #NVIDIA#
Show more
World models are moving beyond offline generation towards interactive, real-time experiences. Introducing ⚡FlashDreams⚡: an open-source high-performance inference and serving library built for autoregressive world models: 🔥 Up to 3.10× faster LingBot-World inference 🔥 Up to 2.12× faster Self-Forcing inference 🔥 Up to 1.40× faster Wan2.1 inference 🔥 8 integrated models 🔥 Multi-GPU, streaming, low-latency serving 🔥 Agentic skills that teach you how to use it FlashDreams is designed for a new generation of AI systems that continuously evolve over time while responding to user interactions. It powers applications across robotics, autonomous vehicle simulation, gaming, and virtual worlds. Github: Docs: Research page: Join the #flashdreams# Discord channel at FlashDreams is also the runtime backbone behind NVIDIA OmniDreams ( 1/n #AI# #WorldModels# #FastInference# #PhysicalAI# #OpenSource# #NVIDIA#
Show more
The latent-vs-pixel debate misses the point. GPT Image 2 shows what users notice: pixel-level fidelity. Latent models show what scales: compact semantic structure. We connect them by replacing VAE/RAE decoders with a Pixel Diffusion Decoder. Code and Model available: 🧵(1/N)
Show more
Our paper has been accepted to #ICML2026# as a Spotlight (top 2.2%) 🥳 We present MOTIVE, a scalable, motion-centric data-attribution framework for video generation that identifies which training clips improve or degrade temporal dynamics, enabling curation and more.
Show more
New #NVIDIA# Paper We introduce Motive, a motion-centric, gradient-based data attribution method that traces which training videos help or hurt video generation. By isolating temporal dynamics from static appearance, Motive identifies which training videos shape motion in video generation. 🔗 1/10
Show more
0
12
636
125
Forward to community
🚀We just released Asset Harvester, an image-to-3D model and end-to-end pipeline that extracts real object assets from autonomous driving videos! 🌐 Website: 💻 Code: [1/5] #AssetHarvester# #AVSimulation# #WorldModel# #AutonomousDriving#
Show more
0
30
785
128
Forward to community
We open-sourced the code and model for UniRelight! 🎉 Given an input video and a target lighting configuration, our method jointly predicts a relit video and its corresponding albedo. Code: Model:
Show more
A canonical open-source data platform for multi-sensor neural 3D reconstruction is here. Introducing NVIDIA NCore — a unified data format and APIs for cameras, lidars, radars, poses, calibrations, and labels, built for AV, robotics, and physical AI. 🔗
Show more
Feed-forward 3D reconstruction should not be limited to predicting one Gaussian per pixel. We introduce TokenGS, which uses learnable tokens to decouple the 3D Gaussian prediction from the image resolution and the number of input views. #CVPR2026Highlight# [1/6]
Show more
We scaled up Lyra to generate explorable 3D worlds! 🚀 Introducing Lyra 2.0 — turning a single image into a 3D world you can walk through, look back, and even drop a robot into 🤖 Code and Model available today! 🌐 Website: (1/N)
Show more
0
27
861
115
Forward to community