KV caching is the fundamental optimization underpinning autoregressive LLM inference.
Transformer layers store keys and values from earlier tokens, then reuse them as new tokens are generated.
This avoids recomputation, but at the cost of memory capacity and bandwidth.