๐ New blog: Pushing the Limits of Serving DeepSeek-V4-Pro
DeepSeek-V4-Pro (1.6T MoE) on H20 reaches 271 output tokens/s at batch size 1, just 1.42ร off B300 on hardware with no native FP4 Tensor Cores.
Together with
@ant_oss, we built a scenario-specific serving stack on SGLang:
- 74.8%โ78.0% peak TPOT reduction at batch size 1 from optimized DSpark
- 1M-token prefill in 43.7s, 36.5% geomean prefill throughput gain
- 10.14ร full-token KV capacity from Humming MXFP4AFP8 + Online C128
- 2.20ร per-GPU decode throughput at 4K (319.9 to 703.2 tok/s/GPU)