@Alibaba_Qwen Qwen3.8-Flash-Next on two
@NVIDIAAI DGX Sparks just got better ๐
It inherited some of the improvements i did on the single-spark recipe and a little extra more!
Default recipe ๐
ใปMulti-turn fix: follow-up replies 5x faster (3.25s โ 0.63s), after tool calls 2.55s โ 1.98s
ใปCode up to +27% faster (333 โ 422 tok/s at 8 streams)
ใป~56 tok/s prose, ~76 code (1 stream), up to ~226 prose / ~422 code (8 streams)
ใปOptional bit-exact decoding across both nodes
ใปCached-token reporting in the API
ใปSafer memory default: no more running both nodes at under 1 GiB free
New opt-in lane on vLLM 0.30 (./start-v030.sh) ๐
ใปFollow-ups in 0.28-0.40s, and repeated prompts hit the cache (0.35s vs 2.7s)
ใปSame decode speed, steadier: fixed a vLLM 0.30 default that made speed swing ยฑ15% between runs
ใปOne small patch instead of seven
ใปOptional FP8 KV: 2.6M tokens of context, 3/3 needles at 200k
Tested for stability for a few hours!
I will keep testing this over the next days and probably make the vLLM 0.30 script the default one in the next version of the recipe โ๏ธ