i came across this llama.cpp fork that uses less VRAM and KV memory.
on my GX10, Qwen 3.5-family Bonsai 27B, 60k context, median of three runs:
q8_0:
• 2,006 MiB KV
• 792 tok/s prefill
• 27.2 tok/s decode
KVarN5 + 1,024-token tail:
• 1,349 MiB KV
• 740 tok/s prefill
• 25.6 tok/s decode
The trade: 32.8% less KV memory for roughly 6% lower throughput
might be a solid option for people trying to squeeze in some more tok/s !
repo: