Register and share your invite link to earn from video plays and referrals.

vLLM
@vllm_project
A high-throughput and memory-efficient inference and serving engine for LLMs. Join to discuss together with the community!
Joined March 2024
36 Following    50.2K Followers
🎉 Congrats to @XiaomiMiMo on MiMo-V2.6: day-0 support in vLLM for both sizes. Two MoEs with text, image, video and audio in one checkpoint: Pro is 1.02T total with 42B active, Flash is 309B with 15B active. 1M context, native FP8 weights. Built-in DFlash spec decoding drafts 7 tokens a step. Same architecture class as V2.5, and the serving pieces carry over: hybrid sliding-window plus global attention, the mimo reasoning and tool-call parsers, DFlash spec decoding. Thanks to the @XiaomiMiMo team for opening the weights and the RL stack! 🙌 🔗
Show more
Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:
Show more