At PyTorch Conference North America 2026, Nicolò Lucchesi, Research Engineer at
@MistralAI and vLLM maintainer, will present joint work with
@AWS and
@RedHat on how disaggregated serving in vLLM has evolved to support the latest generation of hybrid models.
This session will cover the evolution of disaggregated serving for modern hybrid architectures, protocol developments aimed at reducing end-to-end latency, and new key-value pinning mechanisms designed for reliability at scale.
Join us in San Jose on October 20-21:
#
PyTorchCon#
@vllm_project