註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Vivek Galatage
@vivekgalatage
browser baker • chromium contributor • enjoys compilers, systems, teaching • director of engineering: ai & systems @visteon • prev founding eng @browsercompany
加入 February 2010
996 正在關注    21.4K 粉絲
GPU architecture | LLM Inference Handbook Add this to your LLM learning resource bundle. "Before writing or tuning GPU kernels, you need a working model of how a GPU runs code. Without it, suggestions like “increase occupancy” or "reduce shared memory bank conflicts" are just a set of rules to memorize. You don't fully understand when they apply and when they don't. This section explains modern GPU architecture at the level needed for kernel work. The details lean toward NVIDIA hardware because CUDA dominates much of the LLM inference ecosystem today. However, the core concepts apply broadly to AMD GPUs and other parallel accelerators as well."
顯示更多
0
4
610
97
轉發到社區