Agent workloads bring long contexts, bursty traffic, and frequent tool calls, creating new demands for AI serving.
Together with
@AMD , we rebuilt the stack for Kimi K2.6 on AMD Instinct™ MI355X, with scheduler-aware multi-tier KV caching.
Up to 3.2× smaller p99 TTFT and 7.7% higher total-token throughput, with no accuracy tax.
Happy to see more Kimi running on AMD chips!
Tech blog: