Register and share your invite link to earn from video plays and referrals.

Mia
@MiaAI_lab
Building with AI & LLMs | Insights, recipes, tools & honest experiments
385 Following    9.6K Followers
Rejoice single DGX Spark owners! 💫 DeepSeek v4 Flash for single DGX Spark BEATS the 2x DGX Spark version on agentic workflows! 🤯🤯 Haven't done full coding tests but for agentic workflows it BEATS the FP8 version running on vLLM running on dual DGX Sparks. I fully expect it to be great at coding too! However, it's much slower, as expected: Aggregate tok/s (1-2-4-6-8-12) 26.7 tok/s → 32.9 → 46.5 → 54.1 → 58.5 → 58.5 tok/s Per-stream progression: 26.7 → 16.5 → 12.1 → 9.5 → 7.7 → 5.2 tok/s That means up to 59 tok/s across 12 concurrent sessions. This is the best model to run on a single DGX Spark! All credits go to @bleysg! I have published a simple start/stop recipe: I still recommend running DeepSeek v4 Flash on 2x DGX Sparks because the speed difference about 3x. Additional details below 👇
Show more
The value of DGX Sparks just went up significantly thanks to a single release from DeepSeek.
Official DeepSeek v4 Flash weights are out in @huggingface 🔥🔥🔥
0
77
2.1K
178
Forward to community
DeepSeek v4 Flash for 2x @NVIDIAAI DGX Sparks has been updated✨ - Single stream recorded ~72 tok/s decode - 6 concurrent up to ~137 tok/s aggregated It's still the best model to run on 2 DGX Sparks. →
Show more
You can now run GLM-5.2 on 3x @NVIDIAAI DGX Sparks at home 🤯 • 248k context • 15-20 tok/s (content-dependent) • MTP for max context / DSpark optional • NVFP4+AQLM hybrid Still a work in progress — will keep improving over time. →
Show more
Still haven’t found a real use case where Grok 4.5 fails at something I ask vs GPT-5.6 Sol, Fable 5, or Kimi K3. All of them are excellent — but Grok 4.5 wins on price and speed by far.
0
153
767
40
Forward to community
I think @PrismML is onto something. Bonsai 27B is already exceptional at tool calling. If their next release is properly trained for coding, it could be a genuine game changer — especially for users on 12-16GB VRAM GPUs and laptops. I'm rooting for them.
Show more
This model can fit in a 12-16gb VRAM consumer GPUs and laptops and it's almost as good as Nvidia's Qwen3.6-27b NVFP4 for agentic workflows, which is insane 🤯 @PrismML's Bonsai-27B 2-bit is performing exceptionally well! Coding tests next. Full eval results link:
Show more
Here's one way to get the most from Grok 4.5 ✨ Switching effort to "low" gives strong results on most tasks, near-zero quality loss, big usage savings — plus it's the fastest.
Show more
0
106
601
44
Forward to community
I'm going to try the new @NVIDIAAI Nemotron-3-Nano-30B-A3B and compare it to Qwen 3.6 35B in agentic workflows.
FYI the best Qwen 3.6 35b nvfp4 to run is the @NVIDIAAI nvfp4. Do not use unsloth nvfp4, it performs worse. Here's how it compares to bf16 👇