注册并分享邀请链接,可获得视频播放与邀请奖励。

SemiAnalysis
@SemiAnalysis_
加入 January 2024
30 正在关注    160.9K 粉丝
NVIDIA LPU supports 3 types of disaggregated inferencing: 1. Rubin Prefill + LPU Decode for the fastest interactivity 2. Rubin Prefill + Rubin Decode Attention + LPU Decode FFN for the middle of the curve 3. Rubin Prefill + Rubin Decode Verification + LPU Drafter for the middle-left of the curve For low interactivity, raw Rubin still takes the win. Looking forward to seeing Rubin + LPU performance curves on open-source agentic benchmarks like AgentX.
显示更多
0
44
478
37
转发到社区