We’re proud to be the Day 0 open-source inference engine partner for
@Alibaba_Qwen 3.8.
To serve this 2.4T-parameter model across multi-node
@NVIDIAAI Blackwell inference, we optimized DP/EP scaling across nodes, delivering 30%+ faster performance than TP16, plus DSpark speculative decoding with single CUDA graph optimization👇