Register and share your invite link to earn from video plays and referrals.

LMSYS Org
@lmsysorg
Large Model Systems Organization: We developed SGLang @sgl_project ( Chatbot Arena (now @arena), and Vicuna!
Joined August 2024
204 Following    17.5K Followers
🚀New blog: Accelerating Long-Context and Agentic Inference with NVFP4 KV Cache 4-bit KV cache is here! NVFP4 KV in SGLang packs ~1.78× more context into GPU memory and speeds up long-context decode by up to 78%. Together with @Alibaba_Qwen and @nvidia, we brought NVFP4 KV cache to Blackwell: - NVFP4 stores KV in just ~56% of FP8's footprint per token -️ Decode throughput jumps +37% / +58% / +78% at 32K / 160K / 1M context - Near-lossless accuracy: matches FP8 on GPQA-Diamond & AIME 2025 (Qwen3.5-397B-A17B) - Higher cache hit rate keeps AgentX throughput scaling where FP8 drops off The whole recipe combines NVFP4 two-level scaling, in-kernel dequantization on decode, and paged KV cache, and you can enable it in SGLang with a single flag --kv-cache-dtype nvfp4
Show more