Tokens/sec lied to me. Same 30 tasks, 4x DGX Spark:
DeepSeek-V4.1-Flash: 61 tok/s, thinks 39k tokens → 671 s
GLM-5.3-Flash NVFP4 (effort high): 37 tok/s, thinks 5k tokens → 324 s
Same answers. GLM done 2.1x sooner. At effort max GLM thinks like DeepSeek and loses.
Agentic coding (6 repo tasks via OpenCode): DeepSeek 172 s, GLM 243 s. Little thinking there, raw speed wins.
Rule: GLM on high for chat, DeepSeek for tool loops. Full numbers, launcher, harness: