註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Vipul Ved Prakash
@vipulved
Co-founder, CEO @togethercompute
加入 April 2008
1.1K 正在關注    8.5K 粉絲
Nice stats from @OpenRouter that illustrate how @togethercompute delivers solid and scaled performance for agentic workloads. We are serving GLM 5.3 and GLM 5.3 Flash @ top decile of TPS, latency, cache rate, and doing it at large volumes, 23% and 30% of all OpenRouter traffic, and OpenRouter is a fraction of our overall API traffic. Running scaled inference for agentic applications is a lot more than maxxing a single metric ... you need to optimize on all dimensions and doing so at scale with reliability.
顯示更多
0
7
59
11
轉發到社區