vLLM
@vllm_project sessions at #
PyTorchCon# NA 2026 span the serving stack, from attention and KV cache management to disaggregated serving and hardware portability across accelerators.
Speakers will cover attention and KV cache systems, disaggregated serving, expert parallelism, and vLLM across TPU, Trainium, Arm, and IBM Spyre.
Featured speakers include:
@RedHat: Lucas Wilkinson, Matthew Bonanni, Zhanqiu Hu,
@rickynds, and Alex Brooks
@amazon: Sunita Nadampalli
@IBM: Or Ozeri, Thomas Parnell, and Dave Grove
@Huawei: Mengqing Cao
@nvidia: Itay Alroy
@Google:
@Rob_Mulla and Qi Zhou
@awscloud: Maen Suleiman
@Meta: Richard Zou, Colin Taylor, and Angela Yi
@BAAIBeijing: Yonghua Lin
@fujitsulabs: Abhishek Jain and N Maajid Khan
@IBMResearch: Antoni Viros i Martin and Avery Blanchard
@MistralAI: Nicolò Lucchesi
@tensormesh:
@this_will_echo
@googlecloud: Bill Jia
Register by September 4 to save:
Read the vLLM session guide: