Register and share your invite link to earn from video plays and referrals.

Bryan Catanzaro
@ctnzr
VP, Applied Deep Learning Research @ NVIDIA
496 Following    28.1K Followers
Love the idea of models that understand how humans interact! And thanks for building with Nemotron Ultra! 💚
.@poolsideai engineers were awesome to collaborate with to launch open models. Excited for what’s to come to open models, and the @nvidia team working on Nemotron.
A sweeping agreement with startup Poolside aims to build an open artificial-intelligence ecosystem in the U.S. that would compete with Chinese heavyweights and American AI giants.
Show more
Building AI is a huge distributed systems challenge. To post-train NVIDIA Nemotron, we bring data generation, inference, reinforcement learning, and evaluation into one loop across CPUs and GPUs. I’ll be speaking about how we do this using Ray during the Ray Summit. August 26 at 9:50 AM PT. 👉
Show more
We had a chance to test drive a preview version of Nemtron 3.5 Lightning at Lila. We tested it in some really challenging RL environments in materials science and it was super easy to train and mastered the task much faster than previous models. Nice job by the Nemotron team!
Show more
A great American open weights MoE model that can run efficiently on your laptop or local hardware like the DGX Spark! You can use the larger Nemotron Ultra on Perplexity!
We post-trained @NVIDIAAI Nemotron 3.5 Lightning on Legal Agent Bench with @trajectorylabs. Here's what we found: 1) Post-training improved agent performance from 0% to 8.3% on held-out LAB tasks, beating both Opus 4.6 and the much larger post-trained Nemotron 3 Ultra. 2) Performance improved across nine practice areas with no regressions. 3) Post-training reduced average model output from 90k to 37k tokens, increasing the model's reward-per-token by 2.4x. Through our collaboration with NVIDIA and Trajectory we’re committed to pushing the frontier of legal intelligence and cost efficiency with open weight models. Deep dive:
Show more
NVIDIA has just released the first Nemotron 3.5 model: Nemotron 3.5 Lightning, a highly efficient small open weights model with performance similar to gpt-oss-120b at around a quarter of the total parameters Nemotron 3.5 Lightning is the successor to @nvidia Nemotron 3 Nano 30B A3B, with 31.6B total and 3.6B active parameters. It retains the same hybrid Mamba-Transformer architecture and small size from Nemotron 3 Nano, but makes substantial gains in intelligence and agentic performance. Key takeaways: ➤ Major intelligence jump: Nemotron 3.5 Lightning scores 24 on the Artificial Analysis Intelligence Index, a +9 point improvement over Nemotron 3 Nano (15). This puts it in line with OpenAI's gpt-oss-120b (24) and only just behind Nemotron 3 Super (26), a model ~4x its size ➤ Optimized for efficiency: Nemotron 3.5 Lightning sits behind the most intelligent small models in its size class such as Qwen3.6 35B A3B (32) and Muse Glimmer (high, 35) - but it is built for a different point on the frontier. In pre-release testing of a DeepInfra endpoint serving the final NVFP4 weights, we measured median output speeds of nearly 670 tokens per second, much faster than those models are served in the market today ➤ Meaningful agentic gains: the largest improvements over Nemotron 3 Nano come on agentic evaluations in GDPval-AA v2 (+334 ELO, moving past gpt-oss-120b and Nemotron 3 Super) and Terminal-Bench v2.1 (24% vs 7%). Combined with its speed and permissive OpenMDW-1.1 license, this positions Lightning as an efficient workhorse model for high-volume agentic deployments ➤ Near-lossless NVFP4 quantization: as with prior Nemotron releases, the model ships in NVFP4 alongside BF16 weights. We measured the NVFP4 variant at 24 on the Intelligence Index and saw minimal degradation compared to the higher-precision weights Key model details: ➤ 1 million token context window, text-only reasoning model ➤ 31.6B total and 3.6B active parameters ➤ Released under the OpenMDW-1.1 license, open for commercial use without material restrictions ➤ The model weights are available now along with serverless inference from providers including @DeepInfra, @FireworksAI_HQ, @friendliai, @CoreWeave, @gmi_cloud, @nebiusai, and @CrusoeAI
Show more
Nemotron 3.5 Lightning: Same architecture as 3.0 Nano, with added speculative decoding, and with the intelligence of 3.0 Super. ⚡️⚡️⚡️
Introducing NVIDIA Nemotron 3.5 Lightning⚡ An open 30B MoE model with 3B active parameters, built for always-on agents to complete high-volume, specialized tasks faster. It delivers up to 4x the output speed of similar-sized models.
Show more
Agent Distillation can be surprisingly data-efficient. We distilled Inkling-Small into Nemotron-3-Nano on incident-diagnosis problems using trajectories from only 19 problems! Very encouraging for post-training task-specific models in data-poor regimes.
Show more
As a supporter of the open weights ecosystem, we're proud to be a post-training partner for NVIDIA Nemotron. We post-train Nemotron models for customer use cases, de-risk mainline RL runs on our AC2 platform and training stack, and contribute aggregate workload statistics for inference benchmarking. This is how open models get better, and we're excited to keep working closely with @NVIDIAAI.
Show more
.@ssankar says Palantir was able to make Nvidia's Nemotron Ultra model "better than frontier": "I literally almost felt gaslit when, within 24 hours of getting Nemotron up with no post-training, this is vanilla Nemotron Ultra, it did better than frontier." "If you just looked at the numbers, you would say, 'It's nowhere near the Frontier. That shouldn't even be possible.'" "But of course, the benchmarks are wrong. I mean, the benchmarks are right for what the benchmark's measuring, but that's not my business. Those are not the tasks my customers had that they were trying to solve."
Show more
Since launching last week, more than 230 companies and organizations from across the tech sector have signed the "Open Weights and American AI Leadership" open letter. We want to thank these partners for standing up and publicly supporting broader access to AI innovation. A special thanks to @nvidia, @a16z, and @PalantirTech for working with @Microsoft on this effort. These signatories understand that America’s AI leadership will not depend on the success of our frontier models alone, but on our ability to build a strong, secure, and open ecosystem that diffuses AI into every sector. We look forward to continuing to work with our partners and with policymakers to build that open ecosystem in a way that benefits American businesses, empowers American workers, and strengthens the American economy.
Show more
0
58
1.1K
175
Forward to community
This is a very troubling development. OpenAI and Anthropic are free to slow down their own AI development efforts all they want. They can cap their compute spend and cut back their own capabilities in various ways. That would be a huge loss for America, but that is their own business. It is absolutely outrageous, however, for America’s two leading labs to ask our government to advocate global “pacing” constraints be imposed on the entire sector. Calling for a global gatekeeper for AI that will cripple our entire nation’s computational capabilities has obvious anti-competitive effects (especially for open source) -- which makes ongoing fears about regulatory capture all the more credible. But that is secondary to the more important point: In the name of addressing one risk, we open ourselves to an even bigger one. A call for mandatory national surrender to a global pacing agreement / body will not constrain China or other actors from advancing, and it will leave America more vulnerable as a result. We stay at the cutting edge because we must. We can find far more reasonable ways to address frontier model safety without restoring to extreme -- and frankly unworkable -- solutions.
Show more
Attackers have frontier AI. Defenders need a frontier AI ecosystem—the best open and closed models, force-multiplied by a global community. During the Hugging Face incident, closed AI blocked essential forensics. An open-weight frontier model helped contain the intrusion. That’s why we created the Open Secure AI Alliance.
Show more
0
2.1K
36.7K
4.3K
Forward to community
We tested Nemotron 3 Ultra on agentic chip design. The task is RTL coding: a model iteratively writes the code that defines a chip's logic, runs it through a simulator, reads the failures, and rewrites. Across nine categories of real design work, Nemotron 3 Ultra averaged a 97.1% pass rate at 6,629 tokens per iteration, best of the open models tested on both.
Show more
worth pointing out the nemotron folks have been doing exceptional open-weight model work for years now, and in general @nvidia has been a major supporter of many key open source efforts that all frontier labs have benefited from (directly or indirectly)
Show more
The most important word here is *ecosystem*. It's not just about having an open-weight model. Open-weight models are a means to an end. To have a truly strong, open ecosystem, we need four critical frontier-level ingredients: open-weight models, open training datasets, open software stacks, and open process knowledge. Few people realize that NVIDIA actually has been pushing beyond open weights by releasing code and datasets for their Nemotron models, which is something open-weight model developers don't do. Marin further opens up the process knowledge - not just how to train one model, but how to iteratively improve and shape a model given particular goals, custom data, and hardware, e.g., how to design scaling laws and evals to guide architecture and data ablations. Open weights, datasets, software, process knowledge: these are the four critical ingredients (renewable resources) that give everyone the ability to most efficiently turn their compute (consumable resources) into the best models according to their needs and values.
Show more
The knowledge that makes AI useful is diffused. It lives with scientists, engineers, clinicians, firms. For AI to benefit from distributed knowledge, it must itself be distributed. Agree with Jensen that this is a future worth building.
Show more
0
328
8.9K
1.1K
Forward to community
Among open-weight models, do the Chinese have the best pre-training? I measured how well LLMs compress fresh, unseen human text (bits-per-byte). Of Western models, only NVIDIA's Nemotron sits squarely in the frontier cluster (Thinking Machines' Inkling is just behind). You can't measure compression directly for closed models like Fable or Sol... commercial APIs block the log probs you'd need. And my attempts to estimate it by sampling were too noisy. One corollary of challenges estimating logprobs is that KL/log-prob distillation against a commercial API is very hard. Sampling can't recover the tails of the probability distribution, and hidden top-p sampling truncates them altogether. That bpb is where it is for frontier Chinese models suggests to me there is quite strong pretraining on raw - not synthetic - text. [Which is not to say there isn't also lots of synthetic data, in both Western and Chinese models.]
Show more
“Open weights strengthen competition and competition is what keeps the benefits of AI broadly shared rather than concentrated in the hands of few”. We believe this strongly at @perplexity_ai and are co-signing this letter, with @nvidia! Thanks to @a16z for putting this together!
Show more