okay enough stonk talk, time for GPU talk $NBIS
we just ran our widest MLPerf Inference round yet and the full rack Nebius system with NVIDIA GB300 NVL72 took first on DeepSeek R1 in both server and offline (about 603k and 690k tokens a second).
simple version of what happened:
1. we got the new chips (including Vera Rubin, the latest NVIDIA iron, and we were one of only two labs to show it in this round)
2. we racked them (full GB300 NVL72, plus the smaller boxes people actually buy).
3. we put them to the test under MLPerf, the public inference benchmark where somebody else holds the stopwatch
result: the full rack took #
1# on DeepSeek R1 for live and batch traffic, and pushed gpt-oss 120B past a million tokens a second.
When we grew from 8 GPUs to 72 (almost 9x!), speed scaled almost 9x too.
One fast GPU is easy, a rack that stays fast when you multiply it is the product.
another day another milestone