GLM 5.2 @ 19.8 tok/sec on 2 DGX Sparks ⚡️
I switched to a dspark draft model with K=2 (model from
@RedHat_AI)
Acceptance is ~68% because the drafter was trained with the full FP8 model...
Next step is to fine tune the drafter so acceptance on this super quant goes up