注册并分享邀请链接,可获得视频播放与邀请奖励。

samsja
@samsja19
leading research at @PrimeIntellect
加入 March 2020
2.7K 正在关注    9.1K 粉丝
we are super bullish on sparse attention + HiSparse and have been working with the vLLM team on it sparse attention reduces the pressure on memory bandwidth by only selecting k for the attention, but it doesn't reduce KV cache memory storage , in high-throughput wide-EP deployment you want to maximize the batch size of decode to use compute as much as possible, but at long sequence decode you quickly run out of VRAM and can't hold enough parallel requests to saturate the compute HiSparse fixes this by offloading the active KV cache to CPU. It keeps an LRU cache on GPU, and since many of the same K tokens are reused every decode you barely notice the offloading, this allows a massive decrease in memory usage and an increase in concurrency This is super important for RL where throughput is key and we want to be as much as possible in a compute-bound regime tldr: lower memory usage, more concurrency, higher inference throughput, faster RL
显示更多
0
11
221
23
转发到社区