๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

SGLang
@sgl_project
Run LLMs fast at any scale ๐Ÿ”— Join our community For AI tech blogs & deep-dives ๐Ÿ‘‰ @lmsysorg
๊ฐ€์ž… May 2025
44 ํŒ”๋กœ์ž‰ ์ค‘    4.1K ํŒฌ
Google Cloud @googlecloud announced its Day 0 support for Kimi K3 with SGLang! They just published the full guide for serving the 2.8T-param MoE with SGLang on GKE, including DSPARK speculative decoding. A model this size is an infrastructure problem before it's a model problem. A4 and A4X VMs gave SGLang the memory bandwidth and interconnect to keep K3 fast under real concurrency. Three ways to deploy: Model Garden, AI Hypercomputer recipes, GKE. Guide in the comments ๐Ÿ‘‡
๋” ๋ณด๊ธฐ