Big update: Among open-weight models, Kimi K3 (Max) is #
1# in the Agent Arena with +9.75% net-improvement, surpassing GLM-5.2 (Max) at +7.12%, and landed the #
1# spot across 5 signals (see below).
Kimi K3 (Max) is also now #
1# in open-weight in the Frontend Code (1682 pts) and Text (1485 pts) Arenas.
Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model.
Congrats to the
@Kimi_Moonshot team for their contribution to the open ecosystem.