nvidia-smi is a good example. in any systems class, you learn pretty quickly that busy does not always mean useful. nvidia defines gpu-util as the % of time where at least one kernel was running. so 95% utilization does not mean the tensor cores were doing 95% of the math they could do.
highly recommend reading the DATE 2024 paper! they looked at TensorFlow workloads running on an A100. gpu-level utilization was high, but average instruction issue rate was below 50%. tensor-core instructions were below 5.2%.
all the newer gpu architecture changes make more sense - TMA handles more of the data movement instead of making threads do it, blackwell added TMEM so the accumulator doesn't take up so much of the register file, matrix multiply can also run asynchronously
tensor cores don't have to wait as much with those changes