Register and share your invite link to earn from video plays and referrals.

alexa griffith
@alexa_griffith_
book enthusiast, senior principal engineer at Red Hat, Inference FDE, @alexasinput podcast host
Joined November 2019
521 Following    504 Followers
llm-d flow control, chapter 2: shared inference under burst pressure. A GPU pool can have spare capacity on average and still run out during a traffic burst. With @_llm_d_ flow control enabled, excess requests are queued until capacity becomes available.
Show more