가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Applied Compute
@appliedcompute
The best AI is built, not bought. Our platform, Applied Compute Agent Cloud, is now in private beta. Book a demo below.
가입 July 2012
18 팔로잉 중    7.5K
At Kimi K3 scale, memory directly determines how many GPUs you need for training. For example, streaming gradients for Adam updates cuts the host memory peak by ~33%. Another bottleneck was SiTU-GLU activations. Naive autograd saves redundant tensors and consumes >100GB of HBM at ~100k tokens/GPU. A custom operator that streams through a fixed-size workspace saves ~90GB of HBM per GPU.
더 보기