๐ Day-0 support for Ling-3.0-flash from
@AntLingAGI is now live in SGLang! A 124B MoE model built for production agents with:
> Hybrid-linear from step 0 of pretraining: KDA + MLA stacked 5:1, 1/64 sparse MoE
> 10,000+ interactive training environments
> New INT4 and MXFP4 variants, running end-to-end on a single NVIDIA DGX Spark via the Spark-adapted SGLang path
โญ๏ธ What makes long agent runs fast: Ling-3.0-flash natively integrates SGLang HiCache + Mooncake hierarchical caching, cutting TTFT by 60% to over 80% on long inputs.
Try it in your agent stack today!