๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Mia
@MiaAI_lab
Building with AI & LLMs | Insights, recipes, tools & honest experiments
๊ฐ€์ž… July 2022
411 ํŒ”๋กœ์ž‰ ์ค‘    36K ํŒฌ
Qwen3.8 Flash for 2x DGX Sparks just got better ๐Ÿ”ฅ - Faster decode on prose and code. - Faster follow-up replies. - Optional vLLM 0.30 path. Get it here:
@Alibaba_Qwen Qwen3.8-Flash-Next on two @NVIDIAAI DGX Sparks just got better ๐Ÿš€ It inherited some of the improvements i did on the single-spark recipe and a little extra more! Default recipe ๐Ÿ‘‡ ใƒปMulti-turn fix: follow-up replies 5x faster (3.25s โ†’ 0.63s), after tool calls 2.55s โ†’ 1.98s ใƒปCode up to +27% faster (333 โ†’ 422 tok/s at 8 streams) ใƒป~56 tok/s prose, ~76 code (1 stream), up to ~226 prose / ~422 code (8 streams) ใƒปOptional bit-exact decoding across both nodes ใƒปCached-token reporting in the API ใƒปSafer memory default: no more running both nodes at under 1 GiB free New opt-in lane on vLLM 0.30 (./start-v030.sh) ๐ŸŒŸ ใƒปFollow-ups in 0.28-0.40s, and repeated prompts hit the cache (0.35s vs 2.7s) ใƒปSame decode speed, steadier: fixed a vLLM 0.30 default that made speed swing ยฑ15% between runs ใƒปOne small patch instead of seven ใƒปOptional FP8 KV: 2.6M tokens of context, 3/3 needles at 200k Tested for stability for a few hours! I will keep testing this over the next days and probably make the vLLM 0.30 script the default one in the next version of the recipe โœŒ๏ธ
๋” ๋ณด๊ธฐ