13 attention mechanisms AI engineers must know:
(bookmark this)
The tricky part about attention is that these techniques are often discussed together, even though they solve very different problems.
Some reduce KV cache size. Some control which tokens can attend to each other. Others make attention cheaper to compute or improve how KV cache is managed during serving.
So, a better way to organize them is by the bottleneck they actually solve.
Let's do that:
đ. đđŠ đĩđ˛đŽđą đđĩđŽđŋđļđģđ´, đđĩđ˛đģ đđŠ đ°đŽđ°đĩđ˛ đđļđđ˛ đļđ đđĩđ˛ đ¯đŧđđđšđ˛đģđ˛đ°đ¸
â MHA (multi-head attention) gives every query head its own key and value heads, providing maximum flexibility but also the largest KV cache.
â MQA (multi-query attention) makes all query heads share a single KV head, dramatically reducing cache size.
â GQA (grouped query attention) groups query heads and gives each group a shared KV head, balancing memory savings with model quality.
â MLA (multi-head latent attention) compresses keys and values into a smaller latent representation before caching them, reducing KV memory even further.
đŽ. đđđđ˛đģđđļđŧđģ đŊđŽđđđ˛đŋđģđ, đđĩđ˛đģ đđĩđŽđ đđĩđ˛ đēđŧđąđ˛đš đ°đŽđģ đđ˛đ˛ đēđŽđđđ˛đŋđ
â Causal attention only allows each token to look backward, which is what autoregressive LLMs use for generation.
â Bidirectional attention lets tokens attend in both directions, which is useful when understanding the entire input at once.
â Sliding Window Attention restricts each token to nearby context instead of attending across the full sequence.
â StreamingLLM keeps a small set of anchor tokens plus recent context, allowing generation to continue with bounded memory.
đ¯. đđŧđēđŊđđđ˛ đ˛đŗđŗđļđ°đļđ˛đģđ°đ, đđĩđ˛đģ đŽđđđ˛đģđđļđŧđģ đļđđđ˛đšđŗ đļđ đ˛đ đŊđ˛đģđđļđđ˛
â FlashAttention computes exact attention in small tiles that fit in fast on-chip memory, reducing expensive memory movement.
â Sparse Attention skips selected token-to-token connections entirely, reducing how much attention needs to be computed.
đ°. đđŠ đđ˛đŋđđļđģđ´ đ˛đŗđŗđļđ°đļđ˛đģđ°đ, đđĩđ˛đģ đŊđŋđŧđąđđ°đđļđŧđģ đđĩđŋđŧđđ´đĩđŊđđ đļđ đđĩđ˛ đ¯đŧđđđšđ˛đģđ˛đ°đ¸
â PagedAttention stores KV cache in blocks allocated on demand, reducing wasted and fragmented GPU memory.
â RadixAttention organizes shared prefixes so KV states can be reused across requests instead of recomputed.
â Prefix Caching similarly reuses KV blocks for repeated prefixes such as system prompts, reducing redundant prefill work.
The important part is that these techniques are complementary.
An LLM can use GQA to shrink its KV cache, FlashAttention to compute attention efficiently, Sliding Window Attention to limit context interactions, and PagedAttention to manage that cache efficiently in production.
Once you organize them by the bottleneck they solve, the attention landscape becomes much easier to reason about.
I wrote a deeper breakdown of how these techniques evolved and the problem each one solves.
The full article is quoted below.
Thanks for reading.
Cheers! :)