註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

wafer
@wafer_ai
Inference that Keeps Getting Better Wafer learns how your workload behaves and continuously optimizes the serving stack for better performance & reliability
加入 June 2025
10 正在關注    11.5K 粉絲
We just released a massive update on our gpu performance engineering resource list AI Performance Engineering v2 this is the most comprehensive resource list for learning gpu and ai perf engineering link in thread 🧵 this version starts with how a single inference request works, then builds through the cuda execution model, roofline, transformer arithmetic, ttft/tpot/goodput, and kernel optimization. it's an opinionated list on what we think is important to deeply understand how to optimize inference systems. also added: - flashattention-4, blackwell tensor memory, and low-precision tensor cores - continuous batching, kv-cache systems, quantization, speculative decoding, and structured decoding - moe serving, collectives, topology, and prefill/decode disaggregation - blackwell ultra, mi350/cdna 4, ironwood, and trainium3 - kernelbench-verified and sol-execbench, plus a separate watchlist for rubin and cdna 5
顯示更多
0
28
592
71
轉發到社區