注册并分享邀请链接,可获得视频播放与邀请奖励。

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
加入 March 2024
36 正在关注    48.4K 粉丝
Monday afternoon we hosted a a vLLM x @googlecloud x @anyscalecompute happy hour in San Francisco 🥂, featuring core maintainers of vLLM and the builders of TPU. There were drinks, food, and a room full of people working on inference. The timing works well because the vLLM team has been working to make TPUs a first-class backend in vLLM, and we see the progress with each PR: Thank you to @googlecloud, @anyscalecompute, and @inferact for co-hosting!
显示更多