Register and share your invite link to earn from video plays and referrals.

Ming-Yu Liu
@liu_mingyu
VP of Cosmos Lab at NVIDIA | IEEE Fellow
526 Following    11.1K Followers
"Data stops being something you collect. It becomes something you compute." At #SIGGRAPH2026#, @liu_mingyu, VP of NVIDIA's Cosmos Lab, laid out why world models are the data engine of physical AI. Plus, he announced Cosmos-Dreams, a neural closed-loop simulator, and showed off a full demo of it in action. You can watch the full keynote here:
Show more
10 million downloads! 🎉 An incredible milestone for NVIDIA Cosmos—and a testament to the team’s vision, hard work, and commitment to advancing physical AI. Proud of what we’ve built together, and even more excited for what developers and researchers will create next. 💚
Show more
10 million downloads and counting for NVIDIA Cosmos models on @huggingface. 🤗 This milestone belongs to the developers and researchers using open world foundation models to build robots, autonomous vehicles and other physical AI systems. Thank you for downloading, experimenting, and building with us. 💚
Show more
10 million downloads and counting for NVIDIA Cosmos models on @huggingface. 🤗 This milestone belongs to the developers and researchers using open world foundation models to build robots, autonomous vehicles and other physical AI systems. Thank you for downloading, experimenting, and building with us. 💚
Show more
Dreamers deserve tools they can actually build with. An API gives you answers. Open weights give you a foundation — something you can fine-tune, tear apart, learn from, and make yours. That's how one model becomes ten thousand things its creators never imagined: a medical assistant in a language the original team didn't speak, a robot policy trained in a garage, a research idea tested overnight instead of never. Closed models scale usage. Open models scale creation.
Show more
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs both frontier closed models and frontier open models.
Show more
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs both frontier closed models and frontier open models.
Show more
0
16.1K
172.2K
29.5K
Forward to community
Robots need world models that can think and act at robot speed. ⚡ Cosmos 3 Edge on Jetson brings multimodal reasoning and action on-device, small enough for the smallest boards and fast enough for real robot control loops. It delivers higher speed and accuracy than similarly sized open models on the same Jetson hardware, making world models practical for robots, not just cloud demos.
Show more
Excited to share: NVIDIA Cosmos 3 Edge is openly available, bringing frontier world models to NVIDIA Jetson Thor and other local devices for physical AI. Read the blog: Learn more about it at today’s SIGGRAPH keynote—in person or via livestream at 3:45 p.m. PT:
Show more
Excited to share: NVIDIA Cosmos 3 Edge is openly available, bringing frontier world models to NVIDIA Jetson Thor and other local devices for physical AI. Read the blog: Learn more about it at today’s SIGGRAPH keynote—in person or via livestream at 3:45 p.m. PT:
Show more
@KenRoth The biggest risk of AI is the concentration of power in a few dominant providers of proprietary AI assistants. The only solution to AI sovereignty is open source foundation models.
0
119
3.1K
452
Forward to community
New work with @nvidia: evaluating robot policies entirely inside a world model. The policy acts, the model imagines the consequences, and the imagined evals predict real-world results. 🧵 real vs world-model rollout side by side📷
Show more
0
28
825
112
Forward to community
New collaboration between NVIDIA and Physical Intelligence: We propose SC3-Eval that evaluates robot policies entirely inside a world model post-trained from Cosmos3-Nano. The policy acts, the model imagines the consequences, and the imagined evals predict real-world results. 🧵
Show more
🎉 Meet vLLM-Omni v0.22.0, a major upgrade for omnimodal world models and production-grade multimodal serving. 🌍 Day-0 @NVIDIAAI Cosmos 3 world models: text, image, audio, video, and action, in and out. 🤖 Robot serving: DreamZero + OpenPI realtime API. 🎙️ Production TTS: Qwen3-TTS, Qwen3-Omni, VoxCPM2 and more. 🎨 Faster image/video/diffusion: Wan 2.2, HunyuanVideo 1.5, LTX-2.3. ⚡ Broader quantization (FP8/INT8, MXFP4/MXFP8, W4A16, ModelOpt) and hardware coverage. 339 commits, 124 contributors, 52 of them new. Thank you all. 🙌 🔗
Show more
Bill Freeman gives us first a list of warm-up bitter lessons. He keeps the bigger ones for later in the talk. #cvpr2026#
Presenting DreamControl at ICRA 2026 today! 💫 @DvijKalaria, @pushkalkatara and Sangkyung Kwak will present our workflow for building whole-body humanoid AI skills — combining diffusion models and RL to train skills without expensive real-world data collection. 17:35 – 17:45 | Session WeBT3 | Lehar 1-4 More about DreamControl: #ICRA2026# #DreamControl# #Humanoid# #PhysicalAI#
Show more
Today we're shipping Nemotron 3 Ultra. A 550B MoE frontier-intelligence open model built for long-running agents. It delivers 5x faster inference and lowers the cost of complex agentic tasks by up to 30% versus other open frontier models.
Show more
0
237
3.6K
480
Forward to community
Built a traffic control system with Hermes Agent using NVIDIA Cosmos 3 and blockchain using Alpenglow Solana protocol. Hermes Agent monitors fleet state and sends control notifications. How it works: Cars stop at an intersection, must get distributed consensus on Solana protocol, then must receive a token, ordered by consensus resolution time, not arrival order before crossing. It simulates a 3×3 city grid. Cosmos 3 video language model (VLM) generates maneuver intents per car or what to do (straight, turn, merge) and how urgently (physics deadline), from text scene descriptions like "dashcam view of a 4-way urban intersection at night, heavy rain…" and returns structured JSON. Alpenglow consensus protocol runs parameterized gossip and voting rounds among all cars approaching an intersection to agree on crossing order. Each intent becomes a transaction that must finalize before the car can proceed. Every car stops at stop sign lines, waits for distributed agreement to finalize, then gets a token in FIFO order by resolution time. The coordination pipeline: Cosmos intent, Alpenglow protocol consensus, get intersection token. A car can't cross until both the intent is generated and consensus has settled. If consensus latency exceeds the physics deadline, the intent "misses" and the car must resolve a safety override.
Show more
Playing with Nvidia Cosmos3 Super models for image and video generation. Here's obligatory Will Smith eating spaghetti. First few renders were pretty boring, so I went all out on an "energetic stuffing face with spaghetti prompt" here. That's some bottomless spaghetti.
Show more
At 4:20pm today, I will give a talk on Cosmos3 in If you are at CVPR and interested about Cosmos3, please come.