Sparkbench local battle:
Qwen3.8-Flash-Next (NVFP4 TP2, spec off) scored 87.0
GLM-5.3-Flash (NVFP4 TP2, DFlash2) at 85.7
This is a deployment comparison, not a universal model ranking, but Qwen was faster, more reliable, and cleaner here.
The most interesting result wasn’t the score:
GLM hit four ~261K-token repetition loops, while Qwen’s NEXTN path had its own collapse and had to be disabled
Local inference moves fast, but serving-stack correctness matters as much as the model.
Full deployment and results:
显示更多