đ New blog: Pushing the Limits of Serving DeepSeek-V4-Pro
DeepSeek-V4-Pro (1.6T MoE) on H20 reaches 271 output tokens/s at batch size 1, just 1.42Ã off B300 on hardware with no native FP4 Tensor Cores.
Together with
@ant_oss, we built a scenario-specific serving stack on SGLang:
- 74.8%â78.0% peak TPOT reduction at batch size 1 from optimized DSpark
- 1M-token prefill in 43.7s, 36.5% geomean prefill throughput gain
- 10.14Ã full-token KV capacity from Humming MXFP4AFP8 + Online C128
- 2.20Ã per-GPU decode throughput at 4K (319.9 to 703.2 tok/s/GPU)