登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Zhuokai Zhao
@zhuokaiz
AI Research Scientist @Meta. Building scalable intelligence. PhD @UChicagoCS.
参加 April 2024
383 フォロー中    5.3K ファン
Got randomly recommended this video from @robertnishihara. Despite being from last year, it's still one of the best at illustrating the unique challenges of LLM inference: 1. Continuous batching (handle variable-length requests dynamically) 2. Prefill-decode disaggregation (separate compute-heavy prefill from memory-bound decode) 3. PagedAttention for KV cache (efficient GPU memory use, less fragmentation) 4. Prefix-aware routing (route shared prefixes to same replicas) 5. MoE sharding (place experts on different GPUs)
もっと見る
Walk with @robertnishihara & I in NYC with 10% charge 🪫 as we talk through 5 key differences between 𝗟𝗟𝗠 𝗶𝗻𝗳𝗲𝗿𝗲𝗻𝗰𝗲 𝗩𝗦 𝗥𝗲𝗴𝘂𝗹𝗮𝗿 𝗶𝗻𝗳𝗲𝗿𝗲𝗻𝗰𝗲 Let’s see how much we can get through before our mic dies! 🤣
もっと見る