Register and share your invite link to earn from video plays and referrals.

Arena.ai
@arena
Where AI meets the real world. We measure and advance the frontier of AI through community-driven evaluation. We’re hiring →
218 Following    214.5K Followers
Exciting news: Grok Imagine Image 2.0 (Low) by @SpaceXAI has landed in the Text-to-Image Arena at #2# (1320 pts)! This release is not available via API, only in their app. Grok Imagine Image 2.0 (Low) is a significant improvement from Grok Imagine Image Quality (#14# -> #2#). By category, it’s #3# in Product, Branding & Commercial Design, and #2# across: - Art - Portraits - Text Rendering - Cartoon, Anime & Fantasy - Photorealistic & Cinematic Imagery In the Image Edit Arena, Grok Imagine Image 2.0 (Low) also debuted at #2# (1439 pts)! See post below. Congrats to the @SpaceXAI team on this release!
Show more
0
52
1.1K
96
Forward to community
Exciting news: Muse Spark 1.2 (xHigh) by @AIatMeta is #4# in the Text Arena (1498 pts), and has reshaped the Pareto frontier! It is priced at $1.25/$4.25 per MToken. Congrats again to the @AIatMeta team on this release!
Show more
Big news: Opus 5 (Max) is now #1# in the Fullstack Code Arena with 1,699 points! The Fullstack Leaderboard shows overall rankings across AI models on full-stack web development tasks: multi-step reasoning, tool use, and end-to-end app generation. Congrats again to the @AnthropicAI team on Opus 5 (Max)!
Show more
Big news: MiniMax-H3 by @MiniMax_AI is now the #1# open model in Video Arena: across both Text-to-Video and Image-to-Video. This is +280pts over the next best open model, hunyuan-video-1.5, and a huge improvement from Hailuo-2.3 at #27# (1199 pts) and Hailuo-02-pro at #28# (1197 pts). In Image-to-Video specifically, it scored 1476 pts, just 2 pts away from Dreamina Seedance-2.0 (1478), making it tied for#1# overall! Learn more about how MiniMax-H3 performed in Text-to-Video Arena in thread.
Show more
0
50
1.2K
131
Forward to community
Qwen3.8-Max by @Alibaba_Qwen has reshaped the cost-performance Pareto frontier in Frontend Code Arena, with pricing of $2 per input MToken and $6 per output MToken. Top models on the Pareto frontier: - Claude-Opus-5 - Kimi-K3 - Qwen3.8-Max - GLM-5.2 - DeepSeek-V4-Flash Congrats to @Alibaba_Qwen on another major milestone!
Show more
Qwen3.8-Max ranks #2# in Vision Arena scoring 1,305. Second only to Claude Fable 5 (High) which has only a 13pt lead.
Big news: Qwen3.8-Max by @Alibaba_Qwen just landed at #4# on the Frontend Code Arena leaderboard with a score of 1,668! With 1,668 points, Qwen3.8-Max is trailing only Claude Opus 5 (Max) with 1,705 pts and Kimi K3 (Max) with 1,676 pts, on par with Claude Opus 5 (High) with 1669 pts. It also ranks high across all domains: #2# in Consumer Product #3# in Brand & Marketing, Reference-based design, Gaming, and Content Creation Tools #4# in Data & Analytics #5# Simulations Dig into the thread for more details and to see how it also performs on real-world tasks in the Text Arena. Congrats to @Alibaba_Qwen on this huge release!
Show more
0
102
2.2K
195
Forward to community
Today we’re launching AutoEval: a new evaluation methodology that ranks models using reward models based on millions of real Arena user preferences. Highlights: - High-quality evaluation signals calibrated on real preference data - Strong alignment with live human evaluations - Evaluations that are several orders of magnitude faster (hours instead of days) - Support for Text, Vision, Image, and Code Arena AutoEval enables us to evaluate newly launched models much faster and share results with the community sooner. We’ll now show AutoEval estimated scores for new models directly on the leaderboard. More details in the thread. 🧵
Show more
Omg, it's finally here - 4 years old we finally have the technology to re-create this - Opus 5 Max
More from Opus 5 and @petergostev. Put Opus 5 to the test with your own prompts in Agent Mode!
Exciting news: @AnthropicAI's Claude Opus 5 (Max) is #2# in Agent Arena, with Opus 5 (High) right behind at #3#, based on over 7K real-world agentic sessions. A strong debut: it slots in just below #1# Fable 5, and ahead of GPT-5.6 Sol (xHigh). Opus 5 Max is #2# with a net-improvement of 11.88%, and is #1# across both Confirmed Success and Praise vs Complaint signals. The default Opus 5 (High) is #3# with net-improvement of 11.73%, and by signal is #3# in Praise vs Complaint and #4# in Confirmed Success. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology.
Show more
Claude Opus 5 vs GPT 5.6 in Agent Arena Highlights: - Opus (High) and Opus (Max) outperform GPT Sol (xHigh), but at higher cost. - Opus (Medium) matches GPT Sol (xHigh) in performance at approximately the same cost. - Fable 5 (High) achieves the most optimal price vs. performance trade off. Real-world cost reflects total task execution, not just per-token pricing, so additional iterations and tool calls increase overall expense.
Show more
Exciting news: @AnthropicAI's Claude Opus 5 (Max) is #2# in Agent Arena, with Opus 5 (High) right behind at #3#, based on over 7K real-world agentic sessions. A strong debut: it slots in just below #1# Fable 5, and ahead of GPT-5.6 Sol (xHigh). Opus 5 Max is #2# with a net-improvement of 11.88%, and is #1# across both Confirmed Success and Praise vs Complaint signals. The default Opus 5 (High) is #3# with net-improvement of 11.73%, and by signal is #3# in Praise vs Complaint and #4# in Confirmed Success. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology.
Show more
Exciting news: @AnthropicAI's Claude Opus 5 (Max) is #2# in Agent Arena, with Opus 5 (High) right behind at #3#, based on over 7K real-world agentic sessions. A strong debut: it slots in just below #1# Fable 5, and ahead of GPT-5.6 Sol (xHigh). Opus 5 Max is #2# with a net-improvement of 11.88%, and is #1# across both Confirmed Success and Praise vs Complaint signals. The default Opus 5 (High) is #3# with net-improvement of 11.73%, and by signal is #3# in Praise vs Complaint and #4# in Confirmed Success. In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks from a global community of users. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology.
Show more
Announced today: Grok 4.6 is coming on August 7. Get ready to see it in @arena next week!
Grok-4.5 from @SpaceXAI debuts at #13# on the new Agent Arena leaderboard, based on 9.8K live agentic sessions. Compared to Grok 4.3, this release is a significant step forward in agentic performance (#29-#>#13#). It makes major gains on Bash Recovery, catching up to Anthropic and OpenAI models, and shows a substantial increase in confirmed task success, making it far more effective in real-world use. Check out rankings by signal below. Congrats to the SpaceXAI team on the strong Grok-4.5 release!
Show more
Code Arena now measures fullstack capabilities! View overall rankings across AI models on full-stack web development tasks: multi-step reasoning, tool use, and end-to-end app generation. - Kimi K3 (Max) takes #1# - GPT 5.6 Sol (xHigh) at #2# - Claude Fable 5 at #3# See more scores at:
Show more
Code Arena just leveled up with fullstack capabilities 🚀 Introducing the new Fullstack Code Arena. We’re moving beyond frontend prototypes to fullstack development complete with databases, API keys, and fast deployments. Build, iterate, and ship real-world software — all in one place. Models now act as agents in the Code Arena, using structured tool calls to plan, execute, and refine in real time with real world tasks. Read more about it in the thread 🧵
Show more
0
112
2.1K
235
Forward to community
In Frontend Code Arena, Kimi K3 (Max) by @Kimi_Moonshot is ranked #1# among open and #1# overall! It’s the #1# open-weight model in all Domains: Brand & Marketing, Reference-Based Design, Data & Analytics, Consumer Product, Gaming, Simulations, and Content Creation Tools. And for overall, #1# in 5 of 7 Domains, landing #2# only in Gaming and Content Creation Tools.
Show more
0
33
841
109
Forward to community
Big update: Among open-weight models, Kimi K3 (Max) is #1# in the Agent Arena with +9.75% net-improvement, surpassing GLM-5.2 (Max) at +7.12%, and landed the #1# spot across 5 signals (see below). Kimi K3 (Max) is also now #1# in open-weight in the Frontend Code (1682 pts) and Text (1485 pts) Arenas. Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Congrats to the @Kimi_Moonshot team for their contribution to the open ecosystem.
Show more
Check out first impressions of Opus 5 with @petergostev now on Arena's YouTube. Leaderboard scores based on real world use coming soon!
At Arena, we are excited to support a healthy American OSS AI ecosystem. Competition and open progress will drive long-term innovation and benefit businesses globally. We are proud signatories of this letter.
Show more
Open-weight models are essential to a healthy AI ecosystem. Together with others across our industry, we are outlining a path for open-weight models to strengthen American competitiveness and expand economic opportunity, while protecting national security.
Show more
Agent & Coding 🔥🔥🔥
Hy3 by Tencent is #5# in Agent Arena for open-weight models (#25# overall)! It also ranks as the #2# open model in the Frontend Code Arena (#16# overall)! In Agent Arena: Hy3 lands at #25# overall (net -2.2%). Hy3 has strengths in tool-use (recovering well from CLI/bash errors, +2.6% and #25#), its biggest weakness is steerability as it struggles to course-correct when users push back, coming in at -7.1% (#30#). Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Below we break down how Hy3 scored across 5 key signals, drawn from tasks submitted by a global community of users. Here’s an overview on the signals: User-satisfaction proxies - Confirmed Success: an explicit "yes that worked" feedback from the user - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? Tool-use proxies - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) - Tool Hallucination: does it call tools that don't exist Congrats to the @TencentHunyuan team on this release!
Show more
Hy3 by Tencent is #5# in Agent Arena for open-weight models (#25# overall)! It also ranks as the #2# open model in the Frontend Code Arena (#16# overall)! In Agent Arena: Hy3 lands at #25# overall (net -2.2%). Hy3 has strengths in tool-use (recovering well from CLI/bash errors, +2.6% and #25#), its biggest weakness is steerability as it struggles to course-correct when users push back, coming in at -7.1% (#30#). Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Below we break down how Hy3 scored across 5 key signals, drawn from tasks submitted by a global community of users. Here’s an overview on the signals: User-satisfaction proxies - Confirmed Success: an explicit "yes that worked" feedback from the user - Praise vs. Complaint: implicit sentiment in users reactions - Steerability: can the model course-correct when you push back? Tool-use proxies - Bash Recovery: how it recovers from CLI errors (primary signal for tool use) - Tool Hallucination: does it call tools that don't exist Congrats to the @TencentHunyuan team on this release!
Show more