注册并分享邀请链接,可获得视频播放与邀请奖励。

Applied Compute
@appliedcompute
The best AI is built, not bought. Our platform, Applied Compute Agent Cloud, is now in private beta. Book a demo below.
加入 July 2012
18 正在关注    7.5K 粉丝
At Kimi K3 scale, memory directly determines how many GPUs you need for training. For example, streaming gradients for Adam updates cuts the host memory peak by ~33%. Another bottleneck was SiTU-GLU activations. Naive autograd saves redundant tensors and consumes >100GB of HBM at ~100k tokens/GPU. A custom operator that streams through a fixed-size workspace saves ~90GB of HBM per GPU.
显示更多