가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

SemiAnalysis
@SemiAnalysis_
가입 January 2024
30 팔로잉 중    160.9K
NVIDIA LPU supports 3 types of disaggregated inferencing: 1. Rubin Prefill + LPU Decode for the fastest interactivity 2. Rubin Prefill + Rubin Decode Attention + LPU Decode FFN for the middle of the curve 3. Rubin Prefill + Rubin Decode Verification + LPU Drafter for the middle-left of the curve For low interactivity, raw Rubin still takes the win. Looking forward to seeing Rubin + LPU performance curves on open-source agentic benchmarks like AgentX.
더 보기