Register and share your invite link to earn from video plays and referrals.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
Joined May 2026
279 Following    412 Followers
For long-context LLM inference, should the KV cache be offloaded to disk or just recomputed on the GPU? There's no universal right answer, and this paper builds a system that decides quantitatively. Title: Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs URL: ๐Ÿ“ Overview The paper proposes py-kvcache, an external KV cache system for vLLM. Using io_uring for async I/O, even a Python implementation pulls near-full SSD read bandwidth of 13.5GB/s. โ— Problem it solves Existing external KV caches like LMCache have no criteria for when to actually use them โ€” they load from disk unconditionally even when a short prefix or a fast GPU makes recomputation cheaper. โš™๏ธ Methodology It introduces "scheduler-aware preloading," which starts disk reads while a request is still waiting to be scheduled, plus a "break-even gate" that rejects a load whenever it wouldn't improve time-to-first-token. ๐Ÿ“Š Results On LongBench multi-document QA it's 6.02-7.43x faster than GPU recomputation and 2.77-3.64x faster than LMCache. On multi-turn SCBench, native vLLM read 3.4TB from disk with completion times over 1200s, while py-kvcache kept disk reads to 85GB and completion time to 480s. ๐Ÿ–ฅ๏ธ Use cases On a high-end H100, requests often don't even clear the break-even point, so skipping external caching is fine โ€” but on a lower-end RTX 4000 Ada, external caching clearly wins, giving concrete hardware-specific guidance. #LLMInference# #vLLM#
Show more