註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Inferact
@inferact
Building the future of inference through @vllm_project
加入 December 2025
5 正在關注    7K 粉絲
Behind this blog is months of our team's work tuning vLLM on agentic workloads and validating on @SemiAnalysis_ AgentX benchmark. We find that open source models optimized for agentic workloads reach up to 130K tokens/GPU-sec, 106× cheaper than Opus 5 API pricing. vLLM is the open source agentic production serving engine. Inferact optimizes vLLM and builds enterprise inference on top of it. 🚀
顯示更多
New blog is out: vLLM x AgentX: Optimizing for Real-World Agentic Serving. Agent traffic stresses every layer of the serving stack at once. This post walks the full-stack work for optimizing vLLM on Agentic workloads, including the architecture, framework, and runtime optimizations, measured on AgentX, @SemiAnalysis_'s public agentic benchmark. 🧵1/6
顯示更多