Tokens/sec lied to me. Same 30 tasks, 4x DGX Spark:
DeepSeek-V4.1-Flash: 61 tok/s, thinks 39k tokens → 671 s
GLM-5.3-Flash NVFP4 (effort high): 37 tok/s, thinks 5k tokens → 324 s
Same answers. GLM done 2.1x sooner. At effort max GLM thinks like DeepSeek and loses.
Agentic coding (6 repo tasks via OpenCode): DeepSeek 172 s, GLM 243 s. Little thinking there, raw speed wins.
Rule: GLM on high for chat, DeepSeek for tool loops. Full numbers, launcher, harness:
Heads up for DGX Spark owners: capping clocks with `nvidia-smi -lgc 0,2200` keeps the GB10 cooler and the clocks stable but it does NOT persist across reboots or driver resets. Wrap it in a systemd timer (OnCalendar=*:0/10) or your "capped" box
quietly drifts back toward 3 GHz.
I've published this DeepSeek 4.1 Flash recipe for 4x DGX Spark (mainly clone of Mia's one with tweaks and experimental libs etc). Prose C1: 45-> 55 tps. C16: 134 -> 277 tps.