We're adding a few Qwen3.8-Flash-Next recipes to the SGLang cookbook, each verified on the hardware:
- 1x RTX PRO 6000, PLE table in system RAM
- 1x DGX Spark, PLE table on local NVMe
- 2x DGX Spark, TP=2
- NVFP4 support for both the RadixArk export and
@NVIDIAAI's ModelOpt export
Cookbook:
Thanks to everyone who has been testing and sharing feedback. We've been following the community reports closely and fixed a number of issues that came with this new architecture.
More performance work is on the way, and we want to be ready for what comes next from the Qwen family!