登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

samsja
@samsja19
leading research at @PrimeIntellect
参加 March 2020
2.7K フォロー中    9.1K ファン
we are super bullish on sparse attention + HiSparse and have been working with the vLLM team on it sparse attention reduces the pressure on memory bandwidth by only selecting k for the attention, but it doesn't reduce KV cache memory storage , in high-throughput wide-EP deployment you want to maximize the batch size of decode to use compute as much as possible, but at long sequence decode you quickly run out of VRAM and can't hold enough parallel requests to saturate the compute HiSparse fixes this by offloading the active KV cache to CPU. It keeps an LRU cache on GPU, and since many of the same K tokens are reused every decode you barely notice the offloading, this allows a massive decrease in memory usage and an increase in concurrency This is super important for RL where throughput is key and we want to be as much as possible in a compute-bound regime tldr: lower memory usage, more concurrency, higher inference throughput, faster RL
もっと見る