Got randomly recommended this video from
@robertnishihara.
Despite being from last year, it's still one of the best at illustrating the unique challenges of LLM inference:
1. Continuous batching (handle variable-length requests dynamically)
2. Prefill-decode disaggregation (separate compute-heavy prefill from memory-bound decode)
3. PagedAttention for KV cache (efficient GPU memory use, less fragmentation)
4. Prefix-aware routing (route shared prefixes to same replicas)
5. MoE sharding (place experts on different GPUs)