가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Zhuokai Zhao
@zhuokaiz
AI Research Scientist @Meta. Building scalable intelligence. PhD @UChicagoCS.
가입 April 2024
383 팔로잉 중    5.3K 팬
Got randomly recommended this video from @robertnishihara. Despite being from last year, it's still one of the best at illustrating the unique challenges of LLM inference: 1. Continuous batching (handle variable-length requests dynamically) 2. Prefill-decode disaggregation (separate compute-heavy prefill from memory-bound decode) 3. PagedAttention for KV cache (efficient GPU memory use, less fragmentation) 4. Prefix-aware routing (route shared prefixes to same replicas) 5. MoE sharding (place experts on different GPUs)
더 보기
Walk with @robertnishihara & I in NYC with 10% charge 🪫 as we talk through 5 key differences between 𝗟𝗟𝗠 𝗶𝗻𝗳𝗲𝗿𝗲𝗻𝗰𝗲 𝗩𝗦 𝗥𝗲𝗴𝘂𝗹𝗮𝗿 𝗶𝗻𝗳𝗲𝗿𝗲𝗻𝗰𝗲 Let’s see how much we can get through before our mic dies! 🤣
더 보기