๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
๊ฐ€์ž… March 2024
36 ํŒ”๋กœ์ž‰ ์ค‘    48.4K ํŒฌ
Monday afternoon we hosted a a vLLM x @googlecloud x @anyscalecompute happy hour in San Francisco ๐Ÿฅ‚, featuring core maintainers of vLLM and the builders of TPU. There were drinks, food, and a room full of people working on inference. The timing works well because the vLLM team has been working to make TPUs a first-class backend in vLLM, and we see the progress with each PR: Thank you to @googlecloud, @anyscalecompute, and @inferact for co-hosting!
๋” ๋ณด๊ธฐ