登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

SemiAnalysis
@SemiAnalysis_
参加 January 2024
30 フォロー中    160.9K ファン
NVIDIA LPU supports 3 types of disaggregated inferencing: 1. Rubin Prefill + LPU Decode for the fastest interactivity 2. Rubin Prefill + Rubin Decode Attention + LPU Decode FFN for the middle of the curve 3. Rubin Prefill + Rubin Decode Verification + LPU Drafter for the middle-left of the curve For low interactivity, raw Rubin still takes the win. Looking forward to seeing Rubin + LPU performance curves on open-source agentic benchmarks like AgentX.
もっと見る