๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Yifan Qiao
@yifandotqiao
MTS @inferact | ex-Postdoc at Sky Computing Lab @UCBerkeley | PhD @UCLA | Building efficient systems for AI
๊ฐ€์ž… January 2023
420 ํŒ”๋กœ์ž‰ ์ค‘    632 ํŒฌ
DeepSeek V4 Pro on @vllm_project can be 106x cheaper than Opus 5 on the same @SemiAnalysis_ AgentX workload. We got here with a stack of optimizations, all open in vLLM, and a blog that covers what works and, perhaps more interestingly, what doesn't. This is why I'm excited about open source. Proud of the @inferact team. Try it on your own agents ๐Ÿš€
๋” ๋ณด๊ธฐ