Inference that Keeps Getting Better
Wafer learns how your workload behaves and continuously optimizes the serving stack for better performance & reliability
We just released a massive update on our gpu performance engineering resource list
AI Performance Engineering v2
this is the most comprehensive resource list for learning gpu and ai perf engineering
link in thread ๐งต
this version starts with how a single inference request works, then builds through the cuda execution model, roofline, transformer arithmetic, ttft/tpot/goodput, and kernel optimization.
it's an opinionated list on what we think is important to deeply understand how to optimize inference systems.
also added:
- flashattention-4, blackwell tensor memory, and low-precision tensor cores
- continuous batching, kv-cache systems, quantization, speculative decoding, and structured decoding
- moe serving, collectives, topology, and prefill/decode disaggregation
- blackwell ultra, mi350/cdna 4, ironwood, and trainium3
- kernelbench-verified and sol-execbench, plus a separate watchlist for rubin and cdna 5