註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

SemiAnalysis
@SemiAnalysis_
加入 January 2024
30 正在關注    160.9K 粉絲
NVIDIA LPU supports 3 types of disaggregated inferencing: 1. Rubin Prefill + LPU Decode for the fastest interactivity 2. Rubin Prefill + Rubin Decode Attention + LPU Decode FFN for the middle of the curve 3. Rubin Prefill + Rubin Decode Verification + LPU Drafter for the middle-left of the curve For low interactivity, raw Rubin still takes the win. Looking forward to seeing Rubin + LPU performance curves on open-source agentic benchmarks like AgentX.
顯示更多
0
44
478
37
轉發到社區