登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Akshay 🚀
@akshay_pachaar
Simplifying LLMs, AI Agents, RAG, and Machine Learning for you! • Co-founder @dailydoseofds_• BITS Pilani • 3 Patents • ex-AI Engineer @ LightningAI
参加 July 2012
501 フォロー中    290K ファン
where does all the VRAM go during LLM inference? (4 ways GPU memory is used) loading the model is only the first part of the memory story. once inference starts, GPU memory gets divided across multiple components, and some of them keep growing as context length, batch size, and concurrency increase. the graphic breaks it into four useful buckets: → model weights are the mostly fixed part. once the model is loaded, their memory footprint stays roughly constant. the biggest lever here is precision. moving from FP16/BF16 to INT8 or INT4 reduces the number of bytes needed to store each parameter. → KV cache grows as generation continues. for every previous token, the model stores key and value tensors so attention can reuse them instead of recomputing the entire sequence. longer contexts mean a larger KV cache, and more concurrent requests mean more active caches sitting in memory. → activations and workspace hold temporary intermediate values needed while running attention, MLP layers, kernels, and other computations. this memory is reused across inference steps, but its size can still change with sequence length, batch size, and the kernels being executed. → runtime overhead comes from everything around the model itself. CUDA kernels, memory allocators, metadata, serving-engine buffers, and other runtime structures all consume some VRAM. it is usually smaller than the other buckets, but it is never zero. this is why “the model fits on the GPU” and “the workload fits on the GPU” are two different statements. a model may load comfortably, then run out of memory when you increase the context window, serve more users simultaneously, or increase the batch size. it also explains why quantization can help beyond simply fitting a larger model. shrinking the weight footprint creates room that can instead be used for larger KV caches, more concurrent requests, or bigger batches. that is the broader GPU lesson too. performance is not just about how much arithmetic a GPU can do. it is also about what data occupies memory, how much of it moves during inference, and how often that data can be reused. i wrote the full breakdown of how GPUs actually work and why memory movement sits at the center of LLM inference performance. the article is quoted below.
もっと見る