Register and share your invite link to earn from video plays and referrals.

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
Joined March 2024
36 Following    48.4K Followers
Monday afternoon we hosted a a vLLM x @googlecloud x @anyscalecompute happy hour in San Francisco 🥂, featuring core maintainers of vLLM and the builders of TPU. There were drinks, food, and a room full of people working on inference. The timing works well because the vLLM team has been working to make TPUs a first-class backend in vLLM, and we see the progress with each PR: Thank you to @googlecloud, @anyscalecompute, and @inferact for co-hosting!
Show more