Register and share your invite link to earn from video plays and referrals.

LMSYS Org
@lmsysorg
Large Model Systems Organization: We developed SGLang @sgl_project ( Chatbot Arena (now @arena), and Vicuna!
202 Following    16.7K Followers
🎉 Day-0 support for Ling-3.0-flash from @AntLingAGI is now live in SGLang! A 124B MoE model built for production agents with: > Hybrid-linear from step 0 of pretraining: KDA + MLA stacked 5:1, 1/64 sparse MoE > 10,000+ interactive training environments > New INT4 and MXFP4 variants, running end-to-end on a single NVIDIA DGX Spark via the Spark-adapted SGLang path ⭐️ What makes long agent runs fast: Ling-3.0-flash natively integrates SGLang HiCache + Mooncake hierarchical caching, cutting TTFT by 60% to over 80% on long inputs. Try it in your agent stack today!
Show more
Inkling-small is out today! With SGLang, you can get 648 tok/s decode with DSpark (simulated acc len=4) and 288 tok/s w/o DSpark, under the same setup (8x @NVIDIAAI B200, TP 8, NVFP4, bs=1). What makes this model different is the size. 276B total with 12B active is a sweet spot for RL, and both LoRA and full-parameter training become well within reach. Miles is ready and verified for multimodal RL on Inkling-small, so you can turn your multimodal data into real capability gains. At ~1/4 the size, Inkling-small matches the bigger version in capability and even wins on some benchmarks. Run Inkling-small with SGLang, and customize it with Miles.
Show more
SGLang now supports DSpark, enabling confidence-driven, variable-length verification for speculative decoding 🎉 DSpark addresses a key bottleneck under load: instead of verifying every draft token, it verifies only where the draft model is confident, so the gains hold even as batch size scales. We heavily optimized variable-length verification in SGLang. Across batch sizes 1 to 256, DSpark gives the best throughput/latency tradeoff on DeepSeek-V4-Flash, ahead of both MTP and non-spec. At high concurrency, dynamic scheduling provides up to ~20% higher throughput compared to a fixed budget, while maintaining high verification quality across workloads. With fused kernels and zero-overhead scheduling, DeepSeek-V4-Pro reaches 383.7 tok/s at B=1 on B300. DSpark is now available in SGLang with support for Qwen3 and DeepSeek-V4. Thanks @deepseek_ai for open-sourcing! Blog with full technical details and commands to run below 👇
Show more
🎉 Congrats on the release of Ring-2.6-1T, a trillion-parameter flagship for complex, real-world tasks. Day-0 support is now live in SGLang! ☑️ Enhanced Agent Execution: stable multi-step, tool-calling & long-horizon workflows ☑️ Reasoning Effort Control: high & xhigh modes to tune depth, speed & cost ☑️ Async RL + IcePop: efficient, stable trillion-parameter RL training Run it now with SGLang!
Show more
🚀 Qwen3.6-27B is here, and we have day 0 support on SGLang ✅ 27B params, beats Qwen3.5-397B-A17B across major coding benchmarks → Agentic coding → Text + multimodal reasoning → Thinking / non-thinking modes Smaller model. Bigger results. Try it on SGLang now 🔥
Show more