Register and share your invite link to earn from video plays and referrals.

Bryan Catanzaro
@ctnzr
VP, Applied Deep Learning Research @ NVIDIA
496 Following    27.3K Followers
Agent Distillation can be surprisingly data-efficient. We distilled Inkling-Small into Nemotron-3-Nano on incident-diagnosis problems using trajectories from only 19 problems! Very encouraging for post-training task-specific models in data-poor regimes.
Show more
As a supporter of the open weights ecosystem, we're proud to be a post-training partner for NVIDIA Nemotron. We post-train Nemotron models for customer use cases, de-risk mainline RL runs on our AC2 platform and training stack, and contribute aggregate workload statistics for inference benchmarking. This is how open models get better, and we're excited to keep working closely with @NVIDIAAI.
Show more
.@ssankar says Palantir was able to make Nvidia's Nemotron Ultra model "better than frontier": "I literally almost felt gaslit when, within 24 hours of getting Nemotron up with no post-training, this is vanilla Nemotron Ultra, it did better than frontier." "If you just looked at the numbers, you would say, 'It's nowhere near the Frontier. That shouldn't even be possible.'" "But of course, the benchmarks are wrong. I mean, the benchmarks are right for what the benchmark's measuring, but that's not my business. Those are not the tasks my customers had that they were trying to solve."
Show more
Since launching last week, more than 230 companies and organizations from across the tech sector have signed the "Open Weights and American AI Leadership" open letter. We want to thank these partners for standing up and publicly supporting broader access to AI innovation. A special thanks to @nvidia, @a16z, and @PalantirTech for working with @Microsoft on this effort. These signatories understand that America’s AI leadership will not depend on the success of our frontier models alone, but on our ability to build a strong, secure, and open ecosystem that diffuses AI into every sector. We look forward to continuing to work with our partners and with policymakers to build that open ecosystem in a way that benefits American businesses, empowers American workers, and strengthens the American economy.
Show more
0
58
1.1K
175
Forward to community
This is a very troubling development. OpenAI and Anthropic are free to slow down their own AI development efforts all they want. They can cap their compute spend and cut back their own capabilities in various ways. That would be a huge loss for America, but that is their own business. It is absolutely outrageous, however, for America’s two leading labs to ask our government to advocate global “pacing” constraints be imposed on the entire sector. Calling for a global gatekeeper for AI that will cripple our entire nation’s computational capabilities has obvious anti-competitive effects (especially for open source) -- which makes ongoing fears about regulatory capture all the more credible. But that is secondary to the more important point: In the name of addressing one risk, we open ourselves to an even bigger one. A call for mandatory national surrender to a global pacing agreement / body will not constrain China or other actors from advancing, and it will leave America more vulnerable as a result. We stay at the cutting edge because we must. We can find far more reasonable ways to address frontier model safety without restoring to extreme -- and frankly unworkable -- solutions.
Show more
Attackers have frontier AI. Defenders need a frontier AI ecosystem—the best open and closed models, force-multiplied by a global community. During the Hugging Face incident, closed AI blocked essential forensics. An open-weight frontier model helped contain the intrusion. That’s why we created the Open Secure AI Alliance.
Show more
0
2.1K
36.7K
4.2K
Forward to community
We tested Nemotron 3 Ultra on agentic chip design. The task is RTL coding: a model iteratively writes the code that defines a chip's logic, runs it through a simulator, reads the failures, and rewrites. Across nine categories of real design work, Nemotron 3 Ultra averaged a 97.1% pass rate at 6,629 tokens per iteration, best of the open models tested on both.
Show more
worth pointing out the nemotron folks have been doing exceptional open-weight model work for years now, and in general @nvidia has been a major supporter of many key open source efforts that all frontier labs have benefited from (directly or indirectly)
Show more
The most important word here is *ecosystem*. It's not just about having an open-weight model. Open-weight models are a means to an end. To have a truly strong, open ecosystem, we need four critical frontier-level ingredients: open-weight models, open training datasets, open software stacks, and open process knowledge. Few people realize that NVIDIA actually has been pushing beyond open weights by releasing code and datasets for their Nemotron models, which is something open-weight model developers don't do. Marin further opens up the process knowledge - not just how to train one model, but how to iteratively improve and shape a model given particular goals, custom data, and hardware, e.g., how to design scaling laws and evals to guide architecture and data ablations. Open weights, datasets, software, process knowledge: these are the four critical ingredients (renewable resources) that give everyone the ability to most efficiently turn their compute (consumable resources) into the best models according to their needs and values.
Show more
The knowledge that makes AI useful is diffused. It lives with scientists, engineers, clinicians, firms. For AI to benefit from distributed knowledge, it must itself be distributed. Agree with Jensen that this is a future worth building.
Show more
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs both frontier closed models and frontier open models.
Show more
0
328
9K
1.1K
Forward to community
Among open-weight models, do the Chinese have the best pre-training? I measured how well LLMs compress fresh, unseen human text (bits-per-byte). Of Western models, only NVIDIA's Nemotron sits squarely in the frontier cluster (Thinking Machines' Inkling is just behind). You can't measure compression directly for closed models like Fable or Sol... commercial APIs block the log probs you'd need. And my attempts to estimate it by sampling were too noisy. One corollary of challenges estimating logprobs is that KL/log-prob distillation against a commercial API is very hard. Sampling can't recover the tails of the probability distribution, and hidden top-p sampling truncates them altogether. That bpb is where it is for frontier Chinese models suggests to me there is quite strong pretraining on raw - not synthetic - text. [Which is not to say there isn't also lots of synthetic data, in both Western and Chinese models.]
Show more
“Open weights strengthen competition and competition is what keeps the benefits of AI broadly shared rather than concentrated in the hands of few”. We believe this strongly at @perplexity_ai and are co-signing this letter, with @nvidia! Thanks to @a16z for putting this together!
Show more
The biggest question facing US AI leadership is whether we will treat AI models as infrastructure. Infrastructure (like the internet, electricity, transportation networks) is best built by lots of organizations and countries working together - in cooperation and in competition. The US knows how to build great infrastructure. Every time we have done it, we have opened new horizons of possibility. I believe it is inevitable that we will do the same with AI, and open models will be at the heart of this new infrastructure. Open models are enabling companies and institutions from brand-new startups to titans of every industry to build their own future, to capitalize on their perspective and ideas to better serve customers and solve problems. Open models make sovereignty possible. Policymakers that want to keep America AI at the forefront of this opportunity will understand how necessary open models are: they are now the critical infrastructure of the AI age.
Show more
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs both frontier closed models and frontier open models.
Show more
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter. AI will transform every industry, power every company, and be built by every country. Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty. The world needs both frontier closed models and frontier open models.
Show more
0
16.2K
173.2K
29.8K
Forward to community
At @NVIDIAAI we continue to push open data, techniques and models forward because we know that every organization needs the freedom to build and deploy AI in their own way. We're now the biggest institutional contributor on HuggingFace and we expect to continue publishing. It's not charity or a science project - we know that when AI grows, NVIDIA's opportunities also grow. More analysis on the state of open source AI here:
Show more
Last year’s IMO results from OpenAI and DeepMind were an amazing achievement and very inspiring to me personally. They showed that solving extremely hard math problems with AI is within our reach. This year, we developed a pipeline to reproduce this result with Nemotron and will share it openly to enable others to do the same. Next year, every AI model developed by anyone will be able to earn an IMO Gold Medal. 5/5
Show more
Congratulations to the students who competed at the International Mathematical Olympiad (IMO) 2026. 👏 We put Nemotron 3 Ultra to the test to take on the same problems in the same time limit, with no internet or external tools. The IMO team graded its solutions 30/42, above the 29-point gold threshold. 🥇
Show more
0
28
1.1K
85
Forward to community
Another great open model! Congrats to the team @poolsideai
Today we are releasing Laguna S 2.1. At 118B total parameters, with 8B active per token, it does the work of models several times its size on agentic coding. It is remarkably persistent across long-horizon tasks. And it is small enough to run on a single NVIDIA DGX Spark. It is far more capable than anything we have created before, and I think it redefines what a model in its weight class can do. Laguna S 2.1 is an important model for Poolside. What it represents is even more important. If, five years ago, I had read a book that said that by 2030 everything economically valuable, scientifically interesting, and personally meaningful would be built on intelligence contracted from three or four companies, I would have called it dystopian science fiction. We are at a fork in the road of what kind of world we can have. I believe intelligence should and will become a commodity. The question is whether that intelligence comes from three companies, or from many people who can build it, own it, and shape it. The open ecosystem will not win by being the best in its own category. No one cares who is king of the open-source kingdom. People want the best intelligence for the task they are trying to do, with the right balance of quality, speed, cost, and control. If we want a different future, open models have to be on par with, or better than, their closed equivalents. Laguna S 2.1 is a meaningful step in that direction: capable enough to compete far above its weight class, efficient enough to run on hardware you can own, and open-weight so anyone can build on it. Open-weighting our models is the contribution we can make today toward a world where intelligence can be built and owned by many. And we will keep doing it. I am very proud of this team’s work. A big shout out to everyone at Poolside who made this possible, from infrastructure and data to architecture, pretraining, post-training, evaluations, and inference. Laguna S 2.1 is available today under the OpenMDW-1.1 license, with weights on Hugging Face and access through OpenRouter and our API. We are building toward a future where the most capable intelligence in the world can be owned and shaped by anyone. Laguna S 2.1 is one step. We are going to keep building until that future exists.
Show more
Open models are more secure.
The asymmetry problem When we started the log analysis, we first used "frontier" models behind commercial APIs. This did **not work**: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were **blocked by the providers' safety guardrails**, which cannot distinguish an incident responder from an attacker. We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure. This had a second benefit: no attacker data, and **none of the credentials it referenced, left our environment**.
Show more
I'm so curious how much better AI will become at this for the next World Cup.
👑Nemotron 3 Ultra (@NVIDIAAI): 67% overall, 80% in knockouts — the sharpest knockout caller in the field. Nailed USA 1-4 Belgium, the only model to call it. 14-pick streak, 3 rare hits.
Show more