For long-context LLM inference, should the KV cache be offloaded to disk or just recomputed on the GPU? There's no universal right answer, and this paper builds a system that decides quantitatively.
Title: Building py-kvcache: A Performance Characterization of External KV Caching for vLLM with NVMe SSDs
URL:
📝 Overview
The paper proposes py-kvcache, an external KV cache system for vLLM. Using io_uring for async I/O, even a Python implementation pulls near-full SSD read bandwidth of 13.5GB/s.
❗ Problem it solves
Existing external KV caches like LMCache have no criteria for when to actually use them — they load from disk unconditionally even when a short prefix or a fast GPU makes recomputation cheaper.
⚙️ Methodology
It introduces "scheduler-aware preloading," which starts disk reads while a request is still waiting to be scheduled, plus a "break-even gate" that rejects a load whenever it wouldn't improve time-to-first-token.
📊 Results
On LongBench multi-document QA it's 6.02-7.43x faster than GPU recomputation and 2.77-3.64x faster than LMCache. On multi-turn SCBench, native vLLM read 3.4TB from disk with completion times over 1200s, while py-kvcache kept disk reads to 85GB and completion time to 480s.
🖥️ Use cases
On a high-end H100, requests often don't even clear the break-even point, so skipping external caching is fine — but on a lower-end RTX 4000 Ada, external caching clearly wins, giving concrete hardware-specific guidance.
#
LLMInference# #
vLLM#