benchmark hygiene
๐ง Your inference engine may be WRONG about how much KV cache you really have.
A developer released an open-source tool called cache-pressure that does something surprisingly useful.
It fills your inference server with known contextsโฆ then checks how much of that context is actually still reusable without recomputation.
It was tested on 2ร DGX Spark + DeepSeek V4 Flash ๐
BEFORE fixes
๐ข Advertised KV cache: 2,023,924 tokens
๐งช Actually retained: 1,052,025
โก๏ธ 51.98%
Basically, half the advertised cache survived real pressure.
Then they fixed deduplication + cache-boundary behavior.
AFTER fixes
๐ข Advertised: 2,047,043 tokens
๐งช Effectively reusable: 3,000,048
โก๏ธ 146.56%No, they didn't magically create extra VRAM. ๐
That >100% number reflects effective reusable context under prefix-cache semantics, not literal physical KV capacity.
The tool has already been tested against:๐ฆ llama.cpp
โก vLLM
๐ NInfer
๐ฅ SGLang
This is the benchmark I want to see more often. How much context is actually still there when the system is under load?
Link to post in ALT.