Visualized tutorials for Kimi K3 model architecture and SGLang optimizations are now live on the SGLang blog.
Highlight
- Kimi K3's new Kimi Delta Attention in detail
- SGL hybrid attention Radix tree
- TP / DP strategies, and more
Check out link for the detailed blog!đ
Visualized tutorials for Kimi K3 model architecture and SGLang optimizations. It is finally out now. đĢĩ
We spent a month polishing this blog to introduce everything you need to know about K3, and how SGLang serves it.
- What's K3 secret sauce to support 1M context window on 2.8T model?
- And how SGL support K3's hybrid attention architect?
Link below đ check it out!