註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Tech2Wild
@Tech2Wild
🎮 Tech, gaming, AI, and everything in between. 🤖 Building with it, not just talking about it. 🔥 From the mind of @ToNYD2WiLD
加入 March 2026
203 正在關注    4.3K 粉絲
Just 4×'d our GLM-5.3-Flash KV pool on 4x DGX Spark 📈 1.26M → 5,033,164 fp8 KV tokens. That's 4.8 concurrent requests each at the FULL 1,048,576-token context. Not 262K. The whole million. 320B/18B MoE at NVFP4, on $16K of desk hardware. How: the "residual-headroom rule" — grow the KV slab until only ~8-10 GB stays free per node (32 GiB KV/rank). The trap we gated against: 38 GiB/rank allocates, boots, answers short prompts... then the first 20K-token prefill OOMs a rank and the whole engine dies. On GB10 "it serves" is not the bar. Every KV bump gets gated behind a real long prefill with the engine verified alive after. 32 passes, 38 doesn't.
顯示更多