登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

wafer
@wafer_ai
Inference that Keeps Getting Better Wafer learns how your workload behaves and continuously optimizes the serving stack for better performance & reliability
参加 June 2025
10 フォロー中    11.5K ファン
We just released a massive update on our gpu performance engineering resource list AI Performance Engineering v2 this is the most comprehensive resource list for learning gpu and ai perf engineering link in thread 🧵 this version starts with how a single inference request works, then builds through the cuda execution model, roofline, transformer arithmetic, ttft/tpot/goodput, and kernel optimization. it's an opinionated list on what we think is important to deeply understand how to optimize inference systems. also added: - flashattention-4, blackwell tensor memory, and low-precision tensor cores - continuous batching, kv-cache systems, quantization, speculative decoding, and structured decoding - moe serving, collectives, topology, and prefill/decode disaggregation - blackwell ultra, mi350/cdna 4, ironwood, and trainium3 - kernelbench-verified and sol-execbench, plus a separate watchlist for rubin and cdna 5
もっと見る