Register and share your invite link to earn from video plays and referrals.

Wësche
@WescheNex1q
Day time artist and night time AI enthusiast. Building & benchmarking frontier LLMs on 4x DGX Spark clusters + Mac. Creator of Vesica Studio. Houston
637 Following    2.4K Followers
No comment here
I tried OpenClaw 2.0, and it’s a total disaster. Peter Steinberger has basically unlimited support from OpenAI, infinite tokens, backing from giants like NVIDIA, tailored model optimizations, and an incredibly talented team. Yet after all this time, OpenClaw still just doesn't work. You have to step in and fix things at almost every turn. The UX is so bad that if I asked my own agent to design the interface, it’d probably do a better job. I’m 100% convinced GPT designed it, which explains the quality. A flaw I’ve dealt with since the early Clawdbot days is still there, and the agent gets stuck in loops, repeats itself, and wanders off-task. It’s honestly bizarre how bad it is. Vision alone isn't enough. You need consistency and reliability to last. I’m definitely sticking with the smarter choice: Hermes Agent.
Show more
GLM-5.3 Flash quant showdown on 2× DGX Spark per model. Bench v6.7.1 — 76 scenarios × 2 repeats. EXL3 TR3 4bpw • TrueScore: 90.9 • Capability: 92.2 • Operational: 87.6 • Median turn: 4.30s • Long-response effective rate: ~35.4 tok/s • 53,347 total output tokens NVFP4 • Raw TrueScore: 78.5 • Capability: 71.6 • Operational: 71.9 • Median turn: 13.42s • Long-response effective rate: ~26.5 tok/s • 996,608 total output tokens EXL3 won instruction following (94.2 vs 48.9), structured output (97.2 vs 12.8), code (98.1 vs 68.5), visuals (92.3 vs 36.3), long context (100 vs 85.3) and robustness (100 vs 86.8). NVFP4 won agentic work (98.5 vs 88.0), planning (100 vs 91.2) and narrowly won safety (87.8 vs 85.6). Tool use tied at 78.4. NVFP4 completed every request with zero transport errors, but emitted 18.7× more output and took 3.12× longer per median turn. Important: this is a deployment comparison, not a pure quant-only test. EXL3 used FP8 KV, CUDA graphs and a configured 1M context ceiling. NVFP4 used Marlin, FP8 E4M3 KV, eager mode and a 262K ceiling. Verdict: EXL3 TR3 4bpw was the much better all-around deployment. NVFP4 was genuinely strong for planning and agentic workflows, but its verbosity and format behavior need fixing before it can compete on practical quality and speed.
Show more
2 DGX Sparks. Qwen3.8 Flash Next NVFP4. 64 users, independent prompts + KV caches. 32,768 output tokens in 78.82s: 415.7 tok/s Video: 1K stress. Usable-context test passed 64K/user at C16, zero leakage. Testing Spark limits, not a production recipe. @NVIDIAAI
Show more
I wonder how it will compare with the deepseek we all love, I’ll run some comparison visuals now
Take two: Mac Studio m4 Max recipe Qwen3.8-Flash-Next update Average 83tok/s • MTP off→MTP3 weighted decode: 36.64→68.31 tok/s (1.86×) • MTP3→MTP6 5,999-token decode: 70.44→83.06 tok/s (+17.9%) • TrueScore: 91.8→91.9 •256k context Deployment:
Show more
Hermes agent has so many commands, this is just some of the main ones I need to print this this in my mouse pad @NousResearch
I love my sparks but for my daily agents in leaning more towards Mac
GLM-5.3 Flash and Qwen3.8 Flash Next got the same uncapped prompt “build an interactive museum cube holding a living ocean that evolves continuously from calm water into a rotating supercell and hurricane, then clears back to calm. No cuts.” GLM: 30,396 reasoning / 52,182 generated tokens Qwen: 52,501 / 92,780 Clear winner is glm for this one, as Qwen’s scene needed display-only repairs to become visible and stable. No geometry, assets, or scene concepts were added. Full disclosure is on the page. Watch + interact:
Show more
Sparkbench local battle: Qwen3.8-Flash-Next (NVFP4 TP2, spec off) scored 87.0 GLM-5.3-Flash (NVFP4 TP2, DFlash2) at 85.7 This is a deployment comparison, not a universal model ranking, but Qwen was faster, more reliable, and cleaner here. The most interesting result wasn’t the score: GLM hit four ~261K-token repetition loops, while Qwen’s NEXTN path had its own collapse and had to be disabled Local inference moves fast, but serving-stack correctness matters as much as the model. Full deployment and results:
Show more
I gave GLM-5.3 Flash and Qwen3.8 Flash Next on challenge: Make one metallic sheet visibly fold into an articulated dragon, without a cut or hidden transition. It must awaken, fly, breathe particle fire, return, and unfold into the sheet. The prompt required procedural geometry, 20,000+ particles, a moonlit Japanese temple, articulated wings and tail, moving shadows, fog, bloom, orbit/zoom controls, camera presets, a live HUD, 60fps targeting, and one self-contained HTML file. GLM: 66,961 reasoning / 81,179 generated tokens Qwen: 117,329 / 153,000 Try it here
Show more
Saw Hermes himself today in Olympia Messenger of the gods
We need to get at least 8 people who bought this together and train them to play soccer 4vs4 I volunteer first
JUST IN: Hugging Face unveils a $400 “Microduck” AI robot that can sing, skate, & learn new behaviors through reinforcement learning.
Local flash battle: GLM-5.3 Flash vs Qwen3.8 Flash Next. GLM rendered clean. Qwen needed two tiny fixes, JS declarations + a GLSL float cast GLM-5.3 Flash: ≈30.4K thinking / ≈38.7K generated Qwen3.8 Flash Next: ≈26.3K / ≈44.8K Both live:
Show more
GLM-5.3-Flash, 320B params, one DGX Spark, 256K context, 33.8 tok/s peak. Released 2 days ago Serving recipe with MTP speculative decode already up, incl. the branch pick that makes it 2x faster:
Show more
GLM-5.3-Flash can now be run locally! ✨ Run 3-bit on 128GB RAM via Unsloth GGUF. GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks. Guide: GGUF:
Show more
Got two spark and one spark glm recipes coming up
GLM-5.3-Flash can now be run locally! ✨ Run 3-bit on 128GB RAM via Unsloth GGUF. GLM-5.3-Flash (ox-alpha) rivals Claude Opus 4.8 on DeepSWE, coding & agentic benchmarks. Guide: GGUF:
Show more
Talk about not been able to keep up, haven’t finished testing the flash version they dropped and now full 5.3 dropping in 25 hours. I’ll say it again, we’re in the inflection point of exponential growth
Show more
More good news: GLM-5.3’s weights will be released tomorrow.
One DGX Spark. Qwen3.6-35B 64 users. 700+ tok/s. 32,768 tokens in 54 seconds. 38W Each user has their own prompt and their own KV cache, and vLLM batches every active stream through the GPU each step. Recipe → @NVIDIAAI
Show more
0
64
1.1K
90
Forward to community