๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Simon Mo
@simon_mo_
๊ฐ€์ž… July 2018
368 ํŒ”๋กœ์ž‰ ์ค‘    4.2K ํŒฌ
๐Ÿซก vLLM is your engine of choice for agentic workload. The pareto curve is easily understood; but very hard to optimize against. The @inferact team did amazing work here, please check it out (and join us!)
๋” ๋ณด๊ธฐ
New blog is out: vLLM x AgentX: Optimizing for Real-World Agentic Serving. Agent traffic stresses every layer of the serving stack at once. This post walks the full-stack work for optimizing vLLM on Agentic workloads, including the architecture, framework, and runtime optimizations, measured on AgentX, @SemiAnalysis_'s public agentic benchmark. ๐Ÿงต1/6
๋” ๋ณด๊ธฐ