Google Cloud
@googlecloud announced its Day 0 support for Kimi K3 with SGLang! They just published the full guide for serving the 2.8T-param MoE with SGLang on GKE, including DSPARK speculative decoding.
A model this size is an infrastructure problem before it's a model problem. A4 and A4X VMs gave SGLang the memory bandwidth and interconnect to keep K3 fast under real concurrency.
Three ways to deploy: Model Garden, AI Hypercomputer recipes, GKE.
Guide in the comments ๐