Pair the DSpark speculator with our GLM-5.3 NVFP4 checkpoint for the best vLLM performance on Blackwell.
NVFP4 gives you a 4-bit target that recovers 95%+ of accuracy across evals. DSpark adds faster decoding on top.
To run both: serve RedHatAI/GLM-5.3-NVFP4 and set --speculative-config to method dspark with the GLM-5.3-speculator.dspark model (8 tokens).
Full config on the card:
顯示更多
🚀 D-Spark for GLM-5.3 beats native MTP on 8×B300: 29% faster single-stream decoding and 16% higher peak throughput. On MRCR, acceptance holds through 1M context, averaging 4.293 accepted tokens in the 524K–1M bucket. Try this out and let us know!
🤗
顯示更多