注册并分享邀请链接,可获得视频播放与邀请奖励。

Vipul Ved Prakash
@vipulved
Co-founder, CEO @togethercompute
加入 April 2008
1.1K 正在关注    8.5K 粉丝
Nice stats from @OpenRouter that illustrate how @togethercompute delivers solid and scaled performance for agentic workloads. We are serving GLM 5.3 and GLM 5.3 Flash @ top decile of TPS, latency, cache rate, and doing it at large volumes, 23% and 30% of all OpenRouter traffic, and OpenRouter is a fraction of our overall API traffic. Running scaled inference for agentic applications is a lot more than maxxing a single metric ... you need to optimize on all dimensions and doing so at scale with reliability.
显示更多
0
7
59
11
转发到社区