๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

LMSYS Org
@lmsysorg
Large Model Systems Organization: We developed SGLang @sgl_project ( Chatbot Arena (now @arena), and Vicuna!
๊ฐ€์ž… August 2024
204 ํŒ”๋กœ์ž‰ ์ค‘    17.5K ํŒฌ
๐Ÿš€ New Blog: Running DeepSeek-V4-Flash and Kimi-K3 on Consumer Hardware with SSD Expert Pack These models are far too large for a typical PC's memory. WiCi AI and the SGLang team built SSD Expert Pack: routed experts stay on an NVMe SSD, and the runtime loads only the experts the router selects into a GPU cache. On one RTX 5090, 32 GB RAM, and a 2 TB SSD ๐Ÿ”ธ DeepSeek-V4-Flash MXFP4: 1.85โ€“1.99 tokens/sec decode ๐Ÿ”ธ Kimi-K3 community Q2_K (text-only): ~0.29 tokens/sec decode Thanks to the WiCi AI team for the collaboration! Read the full blog below ๐Ÿ‘‡
๋” ๋ณด๊ธฐ