註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Woosuk Kwon
@woosuk_k
加入 April 2023
826 正在關注    8.9K 粉絲
KV cache management is arguably one of the hardest problems in LLM inference, and advancing it requires the community to learn from and build on one another’s work. Great to see that happening. Amazing work from @lightseekorg on adopting and evolving vLLM's hybrid memory allocator! The Jenga paper from @ChenZha62999224 is especially relevant today as attention architectures become increasingly complex.
顯示更多
0
6
488
49
轉發到社區