Register and share your invite link to earn from video plays and referrals.

Ru7.ai
@Ru7Longcrypto
御妻老师|言论自由,博您一笑|🦅☝️
Joined May 2022
4.9K Following    42.1K Followers
没想到Kimi K3这么强,国产模型不会要(还是已经)弯道超车了吧
Exciting update: Kimi K3 has landed at #4# on the Agent Arena leaderboard, matching Claude Opus 4.8 and GPT-5.6 Sol. If Kimi K3's weights are released on schedule by July 27, it will become the #1# open-weight model. This release marks a major leap in agentic performance over Kimi K2.7 Code (#23# to #4#). Based on 8K+ live agentic sessions, Kimi K3 leads on confirmed task success rate (#1#). It also posts a strong +20.6% on praise vs. complaint (#3#). It currently lags the field in steerability (#14#) and bash recovery (#17#). Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Here's a primer on the 5 signals: User-satisfaction proxies - Confirmed Success: an explicit "yes that worked" feedback from the user - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? Tool-use proxies - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) - Tool Hallucination: does it call tools that don't exist Below we break down how Kimi K3 scored across the 5 signals, drawn from tasks submitted by a global community of users. Congrats @Kimi_Moonshot on another big milestone!
Show more