Qwen3.8-27B just got stupid fast.
Moved it to ½ RTX PRO 6000 and the prefill/decode speed is absurd compared to L4.
latest optimizations with SGLang completely changed the speed profile.
If you’re running this on RTX PRO 6000 or DGX Spark update the repo right now, It’s worth it.