MiMo V2.6 Flash looks like the sweet spot for local inference. It delivers roughly 90-97% of Pro's practical coding and agentic performance while using just 15B active parameters vs 42B. On 4x DGX Spark, rough estimates are around 50-60 tok/s for Flash vs 20-25 tok/s for a heavily quantized Pro.
Introducing Xiaomi MiMo-V2.6 — Pro & Flash.
Frontier intelligence, all the modalities, built in public.
🔹 Two omnimodal models, advancing through scaled reinforcement learning
🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks
🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models
🔹 Stronger coding, computer use, 3D reasoning and creative capabilities
🔹 Open model weights, technical report, RL environments and training code
Blog: