注册并分享邀请链接,可获得视频播放与邀请奖励。

Red Hat AI
@RedHat_AI
Accelerating AI innovation with open platforms and community. The future of AI is open.
加入 May 2018
2.1K 正在关注    12.6K 粉丝
Now stack it with quantization. Pair the speculator above with our NVFP4 checkpoint: a 4-bit, Blackwell-native target, with DSpark's faster decoding on top. Serve RedHatAI/Qwen3.8-27B-NVFP4 with the DSpark speculative-config (method dspark, 8 tokens).
显示更多
🚀 Our DSpark for Qwen3.8-27B beats native MTP with the same 8 speculative tokens on 4×H200. Up to 52% faster single-stream decoding and 23% higher peak throughput. On 8-needle MRCR, it averages 4.30 accepted tokens beyond 1M context. Model: Demo below, see for yourself!
显示更多