A cache is only worth what it hits.
The community's having a KV cache moment. Here's the corner we work in:
In long-context serving, a cache hit can still leave the GPU waiting for data. When KV lives outside GPU memory, how you bring it back matters. That's the problem we set out to solve. FlexKV restores it layer by layer: earlier layers compute while later layers load. Prefetching starts transfers early, while asynchronous writeback helps overlap cache I/O with inference.
Making those hits faster goes hand in hand with making more of them possible. FlexKV compresses KV losslessly, expands cache capacity with CPU RAM, SSDs, and remote storage, reuses prefixes across the cluster, and routes requests to wherever the cache already lives. It sits under your inference engine, so there’s nothing to rewire.
Works across SGLang, vLLM, TensorRT-LLM, and Dynamo. Up to 70% lower TTFT, +16% QPM.