Register and share your invite link to earn from video plays and referrals.

Zhuokai Zhao
@zhuokaiz
AI Research Scientist @Meta. Building scalable intelligence. PhD @UChicagoCS.
Joined April 2024
383 Following    5.3K Followers
Got randomly recommended this video from @robertnishihara. Despite being from last year, it's still one of the best at illustrating the unique challenges of LLM inference: 1. Continuous batching (handle variable-length requests dynamically) 2. Prefill-decode disaggregation (separate compute-heavy prefill from memory-bound decode) 3. PagedAttention for KV cache (efficient GPU memory use, less fragmentation) 4. Prefix-aware routing (route shared prefixes to same replicas) 5. MoE sharding (place experts on different GPUs)
Show more
Walk with @robertnishihara & I in NYC with 10% charge 🪫 as we talk through 5 key differences between 𝗟𝗟𝗠 𝗶𝗻𝗳𝗲𝗿𝗲𝗻𝗰𝗲 𝗩𝗦 𝗥𝗲𝗴𝘂𝗹𝗮𝗿 𝗶𝗻𝗳𝗲𝗿𝗲𝗻𝗰𝗲 Let’s see how much we can get through before our mic dies! 🤣
Show more