🚀 D-Spark for GLM-5.3 beats native MTP on 8×B300: 29% faster single-stream decoding and 16% higher peak throughput. On MRCR, acceptance holds through 1M context, averaging 4.293 accepted tokens in the 524K–1M bucket. Try this out and let us know!
🤗
显示更多