๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
๊ฐ€์ž… March 2024
36 ํŒ”๋กœ์ž‰ ์ค‘    45.4K ํŒฌ
๐ŸŽ‰ Congrats to @poolsideai on Laguna S 2.1, a new open-weight model built for agentic coding and long-horizon work. ๐Ÿง  118B sparse MoE, only 8B active per token, up to 1M context, thinking + no-thinking modes, OpenMDW-1.1 ๐Ÿ” Built to stay on task across long, multi-step runs: plan, call tools, check its work, recover, keep going ๐Ÿ–ฅ๏ธ The official NVFP4 quant runs locally on a single @NVIDIAAI DGX Spark This model is a scale up of the Laguna XS 2.1 architecture and vLLM runs it out of the box. Keep your existing Laguna serve setup, and the poolside_v1 tool-call and reasoning parsers already work.
๋” ๋ณด๊ธฐ