Register and share your invite link to earn from video plays and referrals.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | 🔔 Follow for AI & Vibe Coding Tips 👇
Joined July 2023
549 Following    11.2K Followers
benchmark hygiene 🧠 Your inference engine may be WRONG about how much KV cache you really have. A developer released an open-source tool called cache-pressure that does something surprisingly useful. It fills your inference server with known contexts… then checks how much of that context is actually still reusable without recomputation. It was tested on 2× DGX Spark + DeepSeek V4 Flash 👇 BEFORE fixes 📢 Advertised KV cache: 2,023,924 tokens 🧪 Actually retained: 1,052,025 ➡️ 51.98% Basically, half the advertised cache survived real pressure. Then they fixed deduplication + cache-boundary behavior. AFTER fixes 📢 Advertised: 2,047,043 tokens 🧪 Effectively reusable: 3,000,048 ➡️ 146.56%No, they didn't magically create extra VRAM. 😁 That >100% number reflects effective reusable context under prefix-cache semantics, not literal physical KV capacity. The tool has already been tested against:🦙 llama.cpp ⚡ vLLM 🚀 NInfer 🔥 SGLang This is the benchmark I want to see more often. How much context is actually still there when the system is under load? Link to post in ALT.
Show more