Register and share your invite link to earn from video plays and referrals.

Lei Li
@_TobiasLee
Ph.D. student @hkunlp2020.
1.1K Following    6.5K Followers
@nrehiew_ yeah let us know if you want to see more metrics :)
Our MOPD from MiMo-V2-Flash has been widely adopted in modern post-training pipelines. Now the paper is out with more details & comparison. Check it out:
Show more
Big week for model releases, and Claw-Eval is updating too. MiMo V2.5 Pro now ranks 3rd, and MiMo V2.5 ranks 5th. Next up: DeepSeek V4? 👉🏻
Kimi K2.6 @Kimi_Moonshot is the new leading open-weights agent model, landing at #4# on Claw-Eval (Pass^3: 62.3%). Key takeaways: - 👑 Best open-source agent, period: Pass^3 of 62.3% is the highest of any open-weights model, within 8 points of frontier Claude Opus 4.6 (70.4%). Pass@3 of 80.9% closes most of the gap to closed models. - 💪Frontier-tier robustness: 94.7 (±0.9) — statistically tied with Claude Sonnet 4.6 (94.6) and Claude Opus 4.6 (94.2). K2.6's agent trajectories no longer collapse under perturbation. The open-source agent frontier just moved. Full Leaderboard:
Show more