注册并分享邀请链接,可获得视频播放与邀请奖励。

SGLang
@sgl_project
Run LLMs fast at any scale 🔗 Join our community For AI tech blogs & deep-dives 👉 @lmsysorg
加入 May 2025
53 正在关注    9.5K 粉丝
We're adding a few Qwen3.8-Flash-Next recipes to the SGLang cookbook, each verified on the hardware: - 1x RTX PRO 6000, PLE table in system RAM - 1x DGX Spark, PLE table on local NVMe - 2x DGX Spark, TP=2 - NVFP4 support for both the RadixArk export and @NVIDIAAI's ModelOpt export Cookbook: Thanks to everyone who has been testing and sharing feedback. We've been following the community reports closely and fixed a number of issues that came with this new architecture. More performance work is on the way, and we want to be ready for what comes next from the Qwen family!
显示更多
0
15
265
31
转发到社区