vLLM (
@vllm_project) sessions at #
PyTorchCon# NA 2026 show how teams are optimizing inference, working on PyTorch release validation and compatibility, and using vLLM in broader serving and application workflows.
Speakers will cover CUTLASS, Triton, and Helion; faster model loading with fastsafetensors; quantization and HiFloat; autonomous kernel bring-up and Intel GPU optimization; PyTorch release validation, enterprise agentic inference, stable ABI, and dynamic shapes; vision-language serving; cloud-to-edge workflows; high-speed storage; a serving stack for a physics-constrained generative model; and contributor onboarding and first-PR workflows.
Featured speakers include:
@nvidia: Michael Goldfarb, Guray Ozen, Aastha Jhunjhunwala, and
@MarkMoyou
@Meta: Andrey Talman,
@janeyx99, and Laith Sakka
@Arm: Kavya Sri Chennoju
@IBM: Takeshi Yoshimura and Nili Guy
@RedHat: Markell Rawls, Joseph Groenenboom, Tyler Michael Smith, Sean McGovern, Christopher Leonard, and Maroon Ayoub
@Google: Ankita Luthra and Trinadh Kotturu
@intel: Xiaogang Gu, Qun Yang, Whitney Tsang, and Artur Fierka
@Huawei: Yun Zhao and Haonan Zhang
@IBMResearch: Burkhard Ringlein
@UMNews:
@arunshar08
Register by September 4 to save:
Read the vLLM session guide: