Inside Nemotron and NVIDIA's AI lab: my conversation with Bryan Catanzaro (
@ctnzr).
@nvidia is a chip company. So why does it put hundreds of researchers on building AI models - and then give them away for free? We go deep into the Nemotron models, what it takes to build a top AI lab, and the future of frontier AI.
01:33 - Is open source AI catching the frontier?
05:29 - Do closed labs blocking distillation slow open source down?
07:42 - Is the US falling behind China?
10:30 - Why companies actually choose open models
12:39 - A "crazy" 2008 bet: machine learning on GPUs
15:33 - Working with Andrew Ng and Dario Amodei at Baidu
17:41 - Coming back to NVIDIA: DLSS and the birth of Megatron
21:55 - The real reason NVIDIA builds its own models
24:28 - Is Moore's Law really dead?
33:37 - The Nemotron family: Nano, Super, Ultra
35:09 - Built for agents: why NVIDIA bets on speed
36:02 - How you train a 550B model in 4 bits
39:25 - Hybrid Mamba-Transformer, explained simply
42:31 - Mixture of experts, and why NVIDIA built NVL72 around it
47:26 - Why a 1-million-token context window matters
49:26 - Multi-token prediction: how the model predicts 5 tokens at once
52:47 - Multi-teacher distillation: teaching one model from many
58:01 - Where reinforcement learning goes next
01:00:16 - Inside NVIDIA's research org: "the mission is the boss"
01:04:03 - How NVIDIA decides who gets the GPUs
01:10:53 - Why NVIDIA still feels entrepreneurial after 33 years
01:12:58 - Why Bryan doesn't believe in the singularity
01:17:50 - The AI backlash
01:19:18 - The controversial case: open AI is safer than closed