๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

SGLang
@sgl_project
Run LLMs fast at any scale ๐Ÿ”— Join our community For AI tech blogs & deep-dives ๐Ÿ‘‰ @lmsysorg
๊ฐ€์ž… May 2025
53 ํŒ”๋กœ์ž‰ ์ค‘    9.5K ํŒฌ
Visualized tutorials for Kimi K3 model architecture and SGLang optimizations are now live on the SGLang blog. Highlight - Kimi K3's new Kimi Delta Attention in detail - SGL hybrid attention Radix tree - TP / DP strategies, and more Check out link for the detailed blog!๐Ÿ‘‡
๋” ๋ณด๊ธฐ
Visualized tutorials for Kimi K3 model architecture and SGLang optimizations. It is finally out now. ๐Ÿซต We spent a month polishing this blog to introduce everything you need to know about K3, and how SGLang serves it. - What's K3 secret sauce to support 1M context window on 2.8T model? - And how SGL support K3's hybrid attention architect? Link below ๐Ÿ‘‡ check it out!
๋” ๋ณด๊ธฐ