Hy3 by Tencent is #
5# in Agent Arena for open-weight models (#
25# overall)! It also ranks as the #
2# open model in the Frontend Code Arena (#
16# overall)!
In Agent Arena: Hy3 lands at #
25# overall (net -2.2%). Hy3 has strengths in tool-use (recovering well from CLI/bash errors, +2.6% and #
25#), its biggest weakness is steerability as it struggles to course-correct when users push back, coming in at -7.1% (#
30#).
Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents.
We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model.
Below we break down how Hy3 scored across 5 key signals, drawn from tasks submitted by a global community of users. Here’s an overview on the signals:
User-satisfaction proxies
- Confirmed Success: an explicit "yes that worked" feedback from the user
- Praise vs. Complaint: implicit sentiment in users reactions
- Steerability: can the model course-correct when you push back?
Tool-use proxies
- Bash Recovery: how it recovers from CLI errors (primary signal for tool use)
- Tool Hallucination: does it call tools that don't exist
Congrats to the
@TencentHunyuan team on this release!