็™ป้Œฒใ—ใฆๆ‹›ๅพ…ใƒชใƒณใ‚ฏใ‚’ๅ…ฑๆœ‰ใ™ใ‚‹ใจใ€ๅ‹•็”ปๅ†็”Ÿๅ ฑ้…ฌใจ็ดนไป‹ๅ ฑ้…ฌใ‚’็ฒๅพ—ใงใใพใ™ใ€‚

Akshay ๐Ÿš€
@akshay_pachaar
Simplifying LLMs, AI Agents, RAG, and Machine Learning for you! โ€ข Co-founder @dailydoseofds_โ€ข BITS Pilani โ€ข 3 Patents โ€ข ex-AI Engineer @ LightningAI
ๅ‚ๅŠ  July 2012
501 ใƒ•ใ‚ฉใƒญใƒผไธญ    290K ใƒ•ใ‚กใƒณ
13 attention mechanisms AI engineers must know: (bookmark this) The tricky part about attention is that these techniques are often discussed together, even though they solve very different problems. Some reduce KV cache size. Some control which tokens can attend to each other. Others make attention cheaper to compute or improve how KV cache is managed during serving. So, a better way to organize them is by the bottleneck they actually solve. Let's do that: ๐Ÿญ. ๐—ž๐—ฉ ๐—ต๐—ฒ๐—ฎ๐—ฑ ๐˜€๐—ต๐—ฎ๐—ฟ๐—ถ๐—ป๐—ด, ๐˜„๐—ต๐—ฒ๐—ป ๐—ž๐—ฉ ๐—ฐ๐—ฎ๐—ฐ๐—ต๐—ฒ ๐˜€๐—ถ๐˜‡๐—ฒ ๐—ถ๐˜€ ๐˜๐—ต๐—ฒ ๐—ฏ๐—ผ๐˜๐˜๐—น๐—ฒ๐—ป๐—ฒ๐—ฐ๐—ธ โ†’ MHA (multi-head attention) gives every query head its own key and value heads, providing maximum flexibility but also the largest KV cache. โ†’ MQA (multi-query attention) makes all query heads share a single KV head, dramatically reducing cache size. โ†’ GQA (grouped query attention) groups query heads and gives each group a shared KV head, balancing memory savings with model quality. โ†’ MLA (multi-head latent attention) compresses keys and values into a smaller latent representation before caching them, reducing KV memory even further. ๐Ÿฎ. ๐—”๐˜๐˜๐—ฒ๐—ป๐˜๐—ถ๐—ผ๐—ป ๐—ฝ๐—ฎ๐˜๐˜๐—ฒ๐—ฟ๐—ป๐˜€, ๐˜„๐—ต๐—ฒ๐—ป ๐˜„๐—ต๐—ฎ๐˜ ๐˜๐—ต๐—ฒ ๐—บ๐—ผ๐—ฑ๐—ฒ๐—น ๐—ฐ๐—ฎ๐—ป ๐˜€๐—ฒ๐—ฒ ๐—บ๐—ฎ๐˜๐˜๐—ฒ๐—ฟ๐˜€ โ†’ Causal attention only allows each token to look backward, which is what autoregressive LLMs use for generation. โ†’ Bidirectional attention lets tokens attend in both directions, which is useful when understanding the entire input at once. โ†’ Sliding Window Attention restricts each token to nearby context instead of attending across the full sequence. โ†’ StreamingLLM keeps a small set of anchor tokens plus recent context, allowing generation to continue with bounded memory. ๐Ÿฏ. ๐—–๐—ผ๐—บ๐—ฝ๐˜‚๐˜๐—ฒ ๐—ฒ๐—ณ๐—ณ๐—ถ๐—ฐ๐—ถ๐—ฒ๐—ป๐—ฐ๐˜†, ๐˜„๐—ต๐—ฒ๐—ป ๐—ฎ๐˜๐˜๐—ฒ๐—ป๐˜๐—ถ๐—ผ๐—ป ๐—ถ๐˜๐˜€๐—ฒ๐—น๐—ณ ๐—ถ๐˜€ ๐—ฒ๐˜…๐—ฝ๐—ฒ๐—ป๐˜€๐—ถ๐˜ƒ๐—ฒ โ†’ FlashAttention computes exact attention in small tiles that fit in fast on-chip memory, reducing expensive memory movement. โ†’ Sparse Attention skips selected token-to-token connections entirely, reducing how much attention needs to be computed. ๐Ÿฐ. ๐—ž๐—ฉ ๐˜€๐—ฒ๐—ฟ๐˜ƒ๐—ถ๐—ป๐—ด ๐—ฒ๐—ณ๐—ณ๐—ถ๐—ฐ๐—ถ๐—ฒ๐—ป๐—ฐ๐˜†, ๐˜„๐—ต๐—ฒ๐—ป ๐—ฝ๐—ฟ๐—ผ๐—ฑ๐˜‚๐—ฐ๐˜๐—ถ๐—ผ๐—ป ๐˜๐—ต๐—ฟ๐—ผ๐˜‚๐—ด๐—ต๐—ฝ๐˜‚๐˜ ๐—ถ๐˜€ ๐˜๐—ต๐—ฒ ๐—ฏ๐—ผ๐˜๐˜๐—น๐—ฒ๐—ป๐—ฒ๐—ฐ๐—ธ โ†’ PagedAttention stores KV cache in blocks allocated on demand, reducing wasted and fragmented GPU memory. โ†’ RadixAttention organizes shared prefixes so KV states can be reused across requests instead of recomputed. โ†’ Prefix Caching similarly reuses KV blocks for repeated prefixes such as system prompts, reducing redundant prefill work. The important part is that these techniques are complementary. An LLM can use GQA to shrink its KV cache, FlashAttention to compute attention efficiently, Sliding Window Attention to limit context interactions, and PagedAttention to manage that cache efficiently in production. Once you organize them by the bottleneck they solve, the attention landscape becomes much easier to reason about. I wrote a deeper breakdown of how these techniques evolved and the problem each one solves. The full article is quoted below. Thanks for reading. Cheers! :)
ใ‚‚ใฃใจ่ฆ‹ใ‚‹