13 attention mechanisms AI engineers must know:
(bookmark this)
The tricky part about attention is that these techniques are often discussed together, even though they solve very different problems.
Some reduce KV cache size. Some control which tokens can attend to each other. Others make attention cheaper to compute or improve how KV cache is managed during serving.
So, a better way to organize them is by the bottleneck they actually solve.
Let's do that:
๐ญ. ๐๐ฉ ๐ต๐ฒ๐ฎ๐ฑ ๐๐ต๐ฎ๐ฟ๐ถ๐ป๐ด, ๐๐ต๐ฒ๐ป ๐๐ฉ ๐ฐ๐ฎ๐ฐ๐ต๐ฒ ๐๐ถ๐๐ฒ ๐ถ๐ ๐๐ต๐ฒ ๐ฏ๐ผ๐๐๐น๐ฒ๐ป๐ฒ๐ฐ๐ธ
โ MHA (multi-head attention) gives every query head its own key and value heads, providing maximum flexibility but also the largest KV cache.
โ MQA (multi-query attention) makes all query heads share a single KV head, dramatically reducing cache size.
โ GQA (grouped query attention) groups query heads and gives each group a shared KV head, balancing memory savings with model quality.
โ MLA (multi-head latent attention) compresses keys and values into a smaller latent representation before caching them, reducing KV memory even further.
๐ฎ. ๐๐๐๐ฒ๐ป๐๐ถ๐ผ๐ป ๐ฝ๐ฎ๐๐๐ฒ๐ฟ๐ป๐, ๐๐ต๐ฒ๐ป ๐๐ต๐ฎ๐ ๐๐ต๐ฒ ๐บ๐ผ๐ฑ๐ฒ๐น ๐ฐ๐ฎ๐ป ๐๐ฒ๐ฒ ๐บ๐ฎ๐๐๐ฒ๐ฟ๐
โ Causal attention only allows each token to look backward, which is what autoregressive LLMs use for generation.
โ Bidirectional attention lets tokens attend in both directions, which is useful when understanding the entire input at once.
โ Sliding Window Attention restricts each token to nearby context instead of attending across the full sequence.
โ StreamingLLM keeps a small set of anchor tokens plus recent context, allowing generation to continue with bounded memory.
๐ฏ. ๐๐ผ๐บ๐ฝ๐๐๐ฒ ๐ฒ๐ณ๐ณ๐ถ๐ฐ๐ถ๐ฒ๐ป๐ฐ๐, ๐๐ต๐ฒ๐ป ๐ฎ๐๐๐ฒ๐ป๐๐ถ๐ผ๐ป ๐ถ๐๐๐ฒ๐น๐ณ ๐ถ๐ ๐ฒ๐ ๐ฝ๐ฒ๐ป๐๐ถ๐๐ฒ
โ FlashAttention computes exact attention in small tiles that fit in fast on-chip memory, reducing expensive memory movement.
โ Sparse Attention skips selected token-to-token connections entirely, reducing how much attention needs to be computed.
๐ฐ. ๐๐ฉ ๐๐ฒ๐ฟ๐๐ถ๐ป๐ด ๐ฒ๐ณ๐ณ๐ถ๐ฐ๐ถ๐ฒ๐ป๐ฐ๐, ๐๐ต๐ฒ๐ป ๐ฝ๐ฟ๐ผ๐ฑ๐๐ฐ๐๐ถ๐ผ๐ป ๐๐ต๐ฟ๐ผ๐๐ด๐ต๐ฝ๐๐ ๐ถ๐ ๐๐ต๐ฒ ๐ฏ๐ผ๐๐๐น๐ฒ๐ป๐ฒ๐ฐ๐ธ
โ PagedAttention stores KV cache in blocks allocated on demand, reducing wasted and fragmented GPU memory.
โ RadixAttention organizes shared prefixes so KV states can be reused across requests instead of recomputed.
โ Prefix Caching similarly reuses KV blocks for repeated prefixes such as system prompts, reducing redundant prefill work.
The important part is that these techniques are complementary.
An LLM can use GQA to shrink its KV cache, FlashAttention to compute attention efficiently, Sliding Window Attention to limit context interactions, and PagedAttention to manage that cache efficiently in production.
Once you organize them by the bottleneck they solve, the attention landscape becomes much easier to reason about.
I wrote a deeper breakdown of how these techniques evolved and the problem each one solves.
The full article is quoted below.
Thanks for reading.
Cheers! :)