20 million ~ H100 GPUs equivalent of compute exists globally.
5% ~ of this training for 6 mo roughly produces a 10 trillion parameter SOTA model.. Astra / Fable level.
The vast majority of compute is used for inference.. pre-training is billions of times less efficient than biology.