Nice stats from
@OpenRouter that illustrate how
@togethercompute delivers solid and scaled performance for agentic workloads.
We are serving GLM 5.3 and GLM 5.3 Flash @ top decile of TPS, latency, cache rate, and doing it at large volumes, 23% and 30% of all OpenRouter traffic, and OpenRouter is a fraction of our overall API traffic.
Running scaled inference for agentic applications is a lot more than maxxing a single metric ... you need to optimize on all dimensions and doing so at scale with reliability.