๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
๊ฐ€์ž… March 2024
36 ํŒ”๋กœ์ž‰ ์ค‘    50.3K ํŒฌ
Thank you to the @SemiAnalysis_ team for the shoutout and for the collaboration on AgentX ๐Ÿ™ Benchmarks are only useful when they measure the workloads people actually run, and AgentX measures the real thing: multi-turn, long-context agent traffic. vLLM is the engine for production agentic workloads. For teams serving tokens at scale, revenue depends on optimized inference over long multi-turn contexts. Our AgentX deep dive blog is coming soon! Stay tuned ๐Ÿ“–
๋” ๋ณด๊ธฐ