註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
加入 March 2024
36 正在關注    45.4K 粉絲
🎉 Congrats to @poolsideai on Laguna S 2.1, a new open-weight model built for agentic coding and long-horizon work. 🧠 118B sparse MoE, only 8B active per token, up to 1M context, thinking + no-thinking modes, OpenMDW-1.1 🔁 Built to stay on task across long, multi-step runs: plan, call tools, check its work, recover, keep going 🖥️ The official NVFP4 quant runs locally on a single @NVIDIAAI DGX Spark This model is a scale up of the Laguna XS 2.1 architecture and vLLM runs it out of the box. Keep your existing Laguna serve setup, and the poolside_v1 tool-call and reasoning parsers already work.
顯示更多
0
6
268
22
轉發到社區