đ Congrats to the teams behind vLLM AFD Plugin (Ascend & vLLM,
@StepFun_ai,
@AntGroup, FastAFD): a new experimental plugin under vllm-project that brings Attention-FFN Disaggregation to MoE serving.
Attention and the expert/FFN path are two very different workloads that normally share one topology. AFD runs them as separate services, so you can scale attention and experts independently. Same vLLM serving surface, no fork.
@NVIDIA GPU and Ascend NPU.