MiMo-V2.6 just dropped, and Xiaomi’s Fuli Luo is already teasing the next architecture: MiMo-V3🔥
The core of it, HySparse2, targets three bottlenecks in agentic inference: prefill cost, KV cache size, and long-context retrieval.
Compared with MiMo-V2.6’s Hybrid SWA architecture at 1M tokens:
• 5.02× lower prefill FLOPs
• 4.5× smaller KV Cache
• Better MRCRv2 and RULER-v2 scores, plus lower AgentPPL and LongPPL
The HySparse2 paper comes from Xiaomi’s LLM-Core team, with Fuli Luo as corresponding author and Team Lead.
Interestingly, its references also include DeepSeek-V2, V3.2, V4, and V4.1-Flash, along with OpenAI’s GPT-4.1 evaluation work and the gpt-oss model card.
顯示更多