Exciting update: Kimi K3 has landed at #
4# on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #
1# open-weight model.
This release marks a major leap in agentic performance over Kimi K2.7 Code (#
23# to #
4#). Based on 8K+ live agentic sessions, Kimi K3 leads on confirmed task success rate (#
1#). It also posts a strong +20.6% on praise vs. complaint (#
3#). It currently lags the field in steerability (#
14#) and bash recovery (#
17#).
Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents.
We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model.
Here's a primer on the 5 signals:
User-satisfaction proxies
- Confirmed Success: an explicit "yes that worked" feedback from the user
- Praise vs. Complaint: implicit sentiment in users reactions
- Steerability: can the model course-correct when you push back?
Tool-use proxies
- Bash Recovery: how it recovers from CLI errors (primary signal for tool use)
- Tool Hallucination: does it call tools that don't exist
Below we break down how Kimi K3 scored across the 5 signals, drawn from tasks submitted by a global community of users.
Congrats
@Kimi_Moonshot on another big milestone!