@Alibaba_Qwen Qwen3.8-Flash-Next on two
@NVIDIAAI DGX Sparks just got better 🚀
It inherited some of the improvements i did on the single-spark recipe and a little extra more!
Default recipe 👇
・Multi-turn fix: follow-up replies 5x faster (3.25s → 0.63s), after tool calls 2.55s → 1.98s
・Code up to +27% faster (333 → 422 tok/s at 8 streams)
・~56 tok/s prose, ~76 code (1 stream), up to ~226 prose / ~422 code (8 streams)
・Optional bit-exact decoding across both nodes
・Cached-token reporting in the API
・Safer memory default: no more running both nodes at under 1 GiB free
New opt-in lane on vLLM 0.30 (./start-v030.sh) 🌟
・Follow-ups in 0.28-0.40s, and repeated prompts hit the cache (0.35s vs 2.7s)
・Same decode speed, steadier: fixed a vLLM 0.30 default that made speed swing ±15% between runs
・One small patch instead of seven
・Optional FP8 KV: 2.6M tokens of context, 3/3 needles at 200k
Tested for stability for a few hours!
I will keep testing this over the next days and probably make the vLLM 0.30 script the default one in the next version of the recipe ✌️