注册并分享邀请链接,可获得视频播放与邀请奖励。

alexa griffith
@alexa_griffith_
book enthusiast, senior principal engineer at Red Hat, Inference FDE, @alexasinput podcast host
加入 November 2019
521 正在关注    504 粉丝
llm-d flow control, chapter 2: shared inference under burst pressure. A GPU pool can have spare capacity on average and still run out during a traffic burst. With @_llm_d_ flow control enabled, excess requests are queued until capacity becomes available.
显示更多