AI clusters have expanded from thousands of GPUs to tens of thousands, and increasingly toward hundreds of thousands of accelerators. Once a training or inference workload spans such a massive number of processors, network latency, bandwidth, radix, topology, power consumption, and reliability directly determine accelerator utilization.
If GPUs spend time waiting for data from other GPUs, memory pools, or remote compute resources, theoretical FLOPS become far less meaningful.
This means that, in the AI era, network utilization increasingly determines compute utilization, while connectivity efficiency increasingly determines the efficiency of intelligence production.