๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

SGLang
@sgl_project
Run LLMs fast at any scale ๐Ÿ”— Join our community For AI tech blogs & deep-dives ๐Ÿ‘‰ @lmsysorg
๊ฐ€์ž… May 2025
53 ํŒ”๋กœ์ž‰ ์ค‘    9.5K ํŒฌ
GLM-5.3-Flash (320B-A18B) from @Zai_org drops today with day-0 support in SGLang. You may know it as ox-alpha from the past few days! ๐Ÿ‘€ It's the first native multimodal model in the GLM-5 series, able to review its own output visually and fix what's wrong. ๐Ÿš€ It outperforms GLM-5.2 at 1/10 the cost. Hybrid sparse plus linear attention enables stable 1M long-context performance at extreme cost efficiency. ๐Ÿ’ป It goes beyond coding into professional work: slides, documents, spreadsheets, and finance research, all handled end to end. Run GLM-5.3-Flash with SGLang, and welcome to a new era of efficient, production-ready open intelligence!
๋” ๋ณด๊ธฐ