๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

LMSYS Org
@lmsysorg
Large Model Systems Organization: We developed SGLang @sgl_project ( Chatbot Arena (now @arena), and Vicuna!
๊ฐ€์ž… August 2024
204 ํŒ”๋กœ์ž‰ ์ค‘    17.5K ํŒฌ
๐Ÿš€ New blog: Pushing the Limits of Serving DeepSeek-V4-Pro DeepSeek-V4-Pro (1.6T MoE) on H20 reaches 271 output tokens/s at batch size 1, just 1.42ร— off B300 on hardware with no native FP4 Tensor Cores. Together with @ant_oss, we built a scenario-specific serving stack on SGLang: - 74.8%โ€“78.0% peak TPOT reduction at batch size 1 from optimized DSpark - 1M-token prefill in 43.7s, 36.5% geomean prefill throughput gain - 10.14ร— full-token KV capacity from Humming MXFP4AFP8 + Online C128 - 2.20ร— per-GPU decode throughput at 4K (319.9 to 703.2 tok/s/GPU)
๋” ๋ณด๊ธฐ