Now stack it with quantization. Pair the speculator above with our NVFP4 checkpoint: a 4-bit, Blackwell-native target, with DSpark's faster decoding on top.
Serve RedHatAI/Qwen3.8-27B-NVFP4 with the DSpark speculative-config (method dspark, 8 tokens).
顯示更多
🚀 Our DSpark for Qwen3.8-27B beats native MTP with the same 8 speculative tokens on 4×H200. Up to 52% faster single-stream decoding and 23% higher peak throughput. On 8-needle MRCR, it averages 4.30 accepted tokens beyond 1M context.
Model:
Demo below, see for yourself!
顯示更多