Register and share your invite link to earn from video plays and referrals.

Inferact
@inferact
Building the future of inference through @vllm_project
Joined December 2025
5 Following    7K Followers
Behind this blog is months of our team's work tuning vLLM on agentic workloads and validating on @SemiAnalysis_ AgentX benchmark. We find that open source models optimized for agentic workloads reach up to 130K tokens/GPU-sec, 106ร— cheaper than Opus 5 API pricing. vLLM is the open source agentic production serving engine. Inferact optimizes vLLM and builds enterprise inference on top of it. ๐Ÿš€
Show more
New blog is out: vLLM x AgentX: Optimizing for Real-World Agentic Serving. Agent traffic stresses every layer of the serving stack at once. This post walks the full-stack work for optimizing vLLM on Agentic workloads, including the architecture, framework, and runtime optimizations, measured on AgentX, @SemiAnalysis_'s public agentic benchmark. ๐Ÿงต1/6
Show more