๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

David Hendrickson
@TeksEdge
CEO & Founder | PhD | Startup Advisor | @Columbia | Author Generative Software Engineering | ๐Ÿ”” Follow for AI & Vibe Coding Tips ๐Ÿ‘‡
๊ฐ€์ž… July 2023
549 ํŒ”๋กœ์ž‰ ์ค‘    11.2K ํŒฌ
benchmark hygiene ๐Ÿง  Your inference engine may be WRONG about how much KV cache you really have. A developer released an open-source tool called cache-pressure that does something surprisingly useful. It fills your inference server with known contextsโ€ฆ then checks how much of that context is actually still reusable without recomputation. It was tested on 2ร— DGX Spark + DeepSeek V4 Flash ๐Ÿ‘‡ BEFORE fixes ๐Ÿ“ข Advertised KV cache: 2,023,924 tokens ๐Ÿงช Actually retained: 1,052,025 โžก๏ธ 51.98% Basically, half the advertised cache survived real pressure. Then they fixed deduplication + cache-boundary behavior. AFTER fixes ๐Ÿ“ข Advertised: 2,047,043 tokens ๐Ÿงช Effectively reusable: 3,000,048 โžก๏ธ 146.56%No, they didn't magically create extra VRAM. ๐Ÿ˜ That >100% number reflects effective reusable context under prefix-cache semantics, not literal physical KV capacity. The tool has already been tested against:๐Ÿฆ™ llama.cpp โšก vLLM ๐Ÿš€ NInfer ๐Ÿ”ฅ SGLang This is the benchmark I want to see more often. How much context is actually still there when the system is under load? Link to post in ALT.
๋” ๋ณด๊ธฐ