๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
๊ฐ€์ž… March 2024
36 ํŒ”๋กœ์ž‰ ์ค‘    50.2K ํŒฌ
๐ŸŽ‰ Congrats to @XiaomiMiMo on MiMo-V2.6: day-0 support in vLLM for both sizes. Two MoEs with text, image, video and audio in one checkpoint: Pro is 1.02T total with 42B active, Flash is 309B with 15B active. 1M context, native FP8 weights. Built-in DFlash spec decoding drafts 7 tokens a step. Same architecture class as V2.5, and the serving pieces carry over: hybrid sliding-window plus global attention, the mimo reasoning and tool-call parsers, DFlash spec decoding. Thanks to the @XiaomiMiMo team for opening the weights and the RL stack! ๐Ÿ™Œ ๐Ÿ”—
๋” ๋ณด๊ธฐ
Introducing Xiaomi MiMo-V2.6 โ€” Pro & Flash. Frontier intelligence, all the modalities, built in public. ๐Ÿ”น Two omnimodal models, advancing through scaled reinforcement learning ๐Ÿ”น Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks ๐Ÿ”น Pro scores 46 on the Artificial Analysis Intelligence Index โ€” the highest among open-source models ๐Ÿ”น Stronger coding, computer use, 3D reasoning and creative capabilities ๐Ÿ”น Open model weights, technical report, RL environments and training code Blog๏ผš
๋” ๋ณด๊ธฐ