Just 4×'d our GLM-5.3-Flash KV pool on 4x DGX Spark 📈
1.26M → 5,033,164 fp8 KV tokens.
That's 4.8 concurrent requests each at the FULL 1,048,576-token context. Not 262K. The whole million. 320B/18B MoE at NVFP4, on $16K of desk hardware.
How: the "residual-headroom rule" — grow the KV slab until only ~8-10 GB stays free per node (32 GiB KV/rank).
The trap we gated against: 38 GiB/rank allocates, boots, answers short prompts... then the first 20K-token prefill OOMs a rank and the whole engine dies. On GB10 "it serves" is not the bar. Every KV bump gets gated behind a real long prefill with the engine verified alive after. 32 passes, 38 doesn't.