๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

wafer
@wafer_ai
Inference that Keeps Getting Better Wafer learns how your workload behaves and continuously optimizes the serving stack for better performance & reliability
๊ฐ€์ž… June 2025
10 ํŒ”๋กœ์ž‰ ์ค‘    11.5K ํŒฌ
We just released a massive update on our gpu performance engineering resource list AI Performance Engineering v2 this is the most comprehensive resource list for learning gpu and ai perf engineering link in thread ๐Ÿงต this version starts with how a single inference request works, then builds through the cuda execution model, roofline, transformer arithmetic, ttft/tpot/goodput, and kernel optimization. it's an opinionated list on what we think is important to deeply understand how to optimize inference systems. also added: - flashattention-4, blackwell tensor memory, and low-precision tensor cores - continuous batching, kv-cache systems, quantization, speculative decoding, and structured decoding - moe serving, collectives, topology, and prefill/decode disaggregation - blackwell ultra, mi350/cdna 4, ironwood, and trainium3 - kernelbench-verified and sol-execbench, plus a separate watchlist for rubin and cdna 5
๋” ๋ณด๊ธฐ