Neil Movva (
@neilmovva) started his career at Nvidia, working on GPUs and kernels, and has an unusually deep understanding of inference, from software to chips to power.
We spend a lot of time on each of those layers, how they connect, and where the important tradeoffs are.
What makes this conversation special is how detailed it is (like a 401-level class), yet Neil makes it remarkably clear and easy to follow.
Today he runs Sail Research, a company building infrastructure for agents to make tokens as cheap as possible.
We discuss:
- Latency versus throughput
- Why there are no bad chips, only bad pricing
- The end of kernel engineering
- Buying chips and power no one else wants
- New chip architectures
- Nvidia lore + his contrarian view of the company
- Open source and the frontier labs
I learned a ton. Enjoy!
TIMESTAMPS
0:00 Intro
0:38 Building a “Token Factory”
4:21 The Future of Background Agents
13:09 Nvidia and the GPU Stack
23:27 Chips, Memory, and Transformers
36:14 The Future of AI Training Data
44:32 Chip Scarcity and Compute Arbitrage
52:44 Reinventing the AI Data Center
59:01 Power and the “Scavenger Strategy”
1:10:10 Open vs. Closed AI