註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Zhuokai Zhao
@zhuokaiz
AI Research Scientist @Meta. Building scalable intelligence. PhD @UChicagoCS.
加入 April 2024
383 正在關注    5.3K 粉絲
Got randomly recommended this video from @robertnishihara. Despite being from last year, it's still one of the best at illustrating the unique challenges of LLM inference: 1. Continuous batching (handle variable-length requests dynamically) 2. Prefill-decode disaggregation (separate compute-heavy prefill from memory-bound decode) 3. PagedAttention for KV cache (efficient GPU memory use, less fragmentation) 4. Prefix-aware routing (route shared prefixes to same replicas) 5. MoE sharding (place experts on different GPUs)
顯示更多
Walk with @robertnishihara & I in NYC with 10% charge 🪫 as we talk through 5 key differences between 𝗟𝗟𝗠 𝗶𝗻𝗳𝗲𝗿𝗲𝗻𝗰𝗲 𝗩𝗦 𝗥𝗲𝗴𝘂𝗹𝗮𝗿 𝗶𝗻𝗳𝗲𝗿𝗲𝗻𝗰𝗲 Let’s see how much we can get through before our mic dies! 🤣
顯示更多