Optimizing large scale inference systems is what we do, so we decided to write down what we know.
Our LLM Inference Handbook is a free reference covering TTFT, TPOT, goodput, continuous batching, chunked prefill, prefix caching, KV cache math, prefill-decode disaggregation, quantization, and more.
It includes 20+ interactive visualizations, is updated continuously, and is open to PRs.