註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
加入 March 2024
36 正在關注    50.3K 粉絲
Thank you to the @SemiAnalysis_ team for the shoutout and for the collaboration on AgentX 🙏 Benchmarks are only useful when they measure the workloads people actually run, and AgentX measures the real thing: multi-turn, long-context agent traffic. vLLM is the engine for production agentic workloads. For teams serving tokens at scale, revenue depends on optimized inference over long multi-turn contexts. Our AgentX deep dive blog is coming soon! Stay tuned 📖
顯示更多