Last one of the day
I was in NVIDIA's Rubin room at Hot Chips.
Same loop as this morning. An agent is 20 to 50 inferences, not one. Then they said no one uses one chip. A 100 MW factory. 2 zettaflops of 4-bit inference. 11 petabytes of memory. 800 petabytes a second.
I called this the Agentic CPU. This is the GPU half. Token revenue, not a faster chip.
They said 30x versus Blackwell Ultra, on real silicon. The slide is 2x versus GB300 until you push interactivity
顯示更多