๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Ant Ling
@AntLingAGI
MoE model series with foundation (Ling), reasoning (Ring) and any-to-any (Ming) from Ant Groupโ€™s AGI initiative, @TheInclusionAI.
๊ฐ€์ž… April 2025
5 ํŒ”๋กœ์ž‰ ์ค‘    13.4K ํŒฌ
Today, weโ€™re introducing the Weight Cache Daemon for SGLang. ๐Ÿš€ On Ling-2.6-1T FP8, it reduced weight loading to ~0.63s, up to ~780ร— faster than disk loading, and cut total engine startup from 8.8 minutes to ~0.53 minutes. Hereโ€™s how it works.
๋” ๋ณด๊ธฐ