Monday afternoon we hosted a a vLLM x
@googlecloud x
@anyscalecompute happy hour in San Francisco 🥂, featuring core maintainers of vLLM and the builders of TPU. There were drinks, food, and a room full of people working on inference.
The timing works well because the vLLM team has been working to make TPUs a first-class backend in vLLM, and we see the progress with each PR:
Thank you to
@googlecloud,
@anyscalecompute, and
@inferact for co-hosting!