llm-d flow control, chapter 2: shared inference under burst pressure.
A GPU pool can have spare capacity on average and still run out during a traffic burst. With
@_llm_d_ flow control enabled, excess requests are queued until capacity becomes available.