At Kimi K3 scale, memory directly determines how many GPUs you need for training.
For example, streaming gradients for Adam updates cuts the host memory peak by ~33%.
Another bottleneck was SiTU-GLU activations. Naive autograd saves redundant tensors and consumes >100GB of HBM at ~100k tokens/GPU. A custom operator that streams through a fixed-size workspace saves ~90GB of HBM per GPU.