登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
参加 March 2024
36 フォロー中    50.3K ファン
🎉 Congrats to the teams behind vLLM AFD Plugin (Ascend & vLLM, @StepFun_ai, @AntGroup, FastAFD): a new experimental plugin under vllm-project that brings Attention-FFN Disaggregation to MoE serving. Attention and the expert/FFN path are two very different workloads that normally share one topology. AFD runs them as separate services, so you can scale attention and experts independently. Same vLLM serving surface, no fork. @NVIDIA GPU and Ascend NPU.
もっと見る