Inference that Keeps Getting Better
Wafer learns how your workload behaves and continuously optimizes the serving stack for better performance & reliability
We just released a massive update on our gpu performance engineering resource list
AI Performance Engineering v2
this is the most comprehensive resource list for learning gpu and ai perf engineering
link in thread 🧵
this version starts with how a single inference request works, then builds through the cuda execution model, roofline, transformer arithmetic, ttft/tpot/goodput, and kernel optimization.
it's an opinionated list on what we think is important to deeply understand how to optimize inference systems.
also added:
- flashattention-4, blackwell tensor memory, and low-precision tensor cores
- continuous batching, kv-cache systems, quantization, speculative decoding, and structured decoding
- moe serving, collectives, topology, and prefill/decode disaggregation
- blackwell ultra, mi350/cdna 4, ironwood, and trainium3
- kernelbench-verified and sol-execbench, plus a separate watchlist for rubin and cdna 5