Red Hat AI just shipped DFlash speculator checkpoints for two of
@NVIDIAAI's most powerful open models:
→ Nemotron Ultra 550B
→ Nemotron Super 120B
On math and reasoning: ~5 out of 7 draft tokens accepted on average. On code (HumanEval): ~3.4 out of 7.
Both checkpoints trained with the open source Speculators library from
@vllm_project. Apache 2.0. Validated on NVIDIA B200.
One flag to enable in vLLM:
--spec-model RedHatAI/NVIDIA-Nemotron-3-Ultra-550B-A55B-speculator.dflash --spec-tokens 7 --spec-method dflash
🔗 Ultra 550B:
🔗 Super 120B: