@googlecloud and Inferact are announcing today a partnership to make TPU a first-class citizen in
@vllm_project.
This partnership puts both teams on one engineering roadmap to bring TPU to the broader open model ecosystem, optimizing vLLM as the agentic production serving engine for TPU:
• Production serving features and optimized kernels
• A native PyTorch path via TorchTPU
• Moving towards day-0 support for frontier model releases
We're also launching a community program: shared TPU capacity for open-source contributors, plus dedicated review and design help from the core vLLM maintainers at Inferact.
Everything this collaboration produces is open source. Read the full announcement: