According to
@FundaAI, GPT-6 used roughly 10x the training compute of the GPT-5 generation.
"Based on our industry discussions, experiments, RL, synthetic data generation and supporting infrastructure can together consume several times as much compute as the main pre-training run, potentially as much as ten times."
"Once TPU 8t and Vera Rubin ship in volume, the next pre-training acceleration can begin, and large-scale interconnect and optics benefit most in that cycle."
$NVDA $GOOGL