登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Arena.ai
@arena
Where AI meets the real world. We measure and advance the frontier of AI through community-driven evaluation. We’re hiring →
参加 March 2023
218 フォロー中    215.6K ファン
Hy3 by Tencent is #5# in Agent Arena for open-weight models (#25# overall)! It also ranks as the #2# open model in the Frontend Code Arena (#16# overall)! In Agent Arena: Hy3 lands at #25# overall (net -2.2%). Hy3 has strengths in tool-use (recovering well from CLI/bash errors, +2.6% and #25#), its biggest weakness is steerability as it struggles to course-correct when users push back, coming in at -7.1% (#30#). Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Below we break down how Hy3 scored across 5 key signals, drawn from tasks submitted by a global community of users. Here’s an overview on the signals: User-satisfaction proxies - Confirmed Success: an explicit "yes that worked" feedback from the user - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? Tool-use proxies - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) - Tool Hallucination: does it call tools that don't exist Congrats to the @TencentHunyuan team on this release!
もっと見る