Register and share your invite link to earn from video plays and referrals.

Arena.ai
@arena
Where AI meets the real world. We measure and advance the frontier of AI through community-driven evaluation. We’re hiring →
224 Following    229.6K Followers
We analyzed how @claudeai’s Opus 5.5 writes compared with Opus 5 across high-reasoning Text Arena outputs. 10 of 12 writing measures moved in a better direction. Opus 5.5 should be easier to read: - Long content words fall from 41.7% to 38.6%, the lowest share of any Claude model we analyzed. - Sentences are 17% shorter on average, dropping from 12.14 to 10.03 words. The tradeoff is length. Answers get 6% wordier, rising from 453 to 481 words on average, making Opus 5.5 give the longest answers across the Opus family. It also sounds less recognizably AI on two familiar tells: - 95% fewer em dashes - 73% fewer semicolons But a new giveaway may be emerging. Hedges and caveats such as “perhaps” and “arguably” rise 97%, from 0.39 to 0.77 per 1,000 words, the highest rate of any Claude model we analyzed. What do you think: does Opus 5.5 read more naturally?
Show more
0
83
1.8K
102
Forward to community
ICYMI: Claude Opus 5.5 (Max) landed #1# in the Code Arena: WebDev with a price 60% lower than the next best model: GPT-6 Astra (Max). In the Arena, all scores are based on live, real-world use by our global community. Check out first impressions with our AI capability expert, @petergostev and let us know what you think.
Show more
Scores for GPT-6 Sol and GPT-6 Luna by @OpenAI are coming soon. Head to Arena now to test them. Your votes on real-world agentic tasks power our leaderboard! In the meantime, @petergostev ran GPT-6 Sol head-to-head against GPT-5.6 Sol under matched conditions: the same prompts, with both models using max reasoning. The comparison examines not just the generations, but also total token usage and wall-clock time, revealing how the models differ in output, token use, and latency. Find the results and Peter’s prompts on our YouTube.
Show more
Qwen-Image-2.1 by @Alibaba_Qwen just landed as the #1# open source model in the Image Edit Arena and Text-to-Image Arena! With 1367 pts in the Image Edit Arena, Qwen-Image-2.1 took the #1# spot among open. It landed #16# overall, just 3 pts from GPT-Image-1.5-high-fidelity at #15#. See the leaderboard for the Text-to-Image arena below. Congrats to the @Alibaba_Qwen team on this contribution to the open source ecosystem!
Show more
Worms Armageddon: Fable 3:0 Astra - highlights from 90 mins playing. Neither played super well, but Astra was particularly suicidal, while Fable did have a few good moments.
"Pi reaches the Pareto frontier on both benchmarks by providing just four tools: read, write, edit, and bash. To understand how harness design affects spending, we examine costs across completed attempts, recorded turn counts, and initial context." Sharing this analysis on the Harness Tax by @MelissaPan and others.
Show more
DeepSeek-V4.1-Flash (Max) is a breakthrough in performance to cost efficiency. With +4.87% net improvement at $0.07 cost per median task, it’s reshaped the Pareto frontier for Agent Arena! Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest median task cost. For comparison, it retains: - 98% of Hy4 preview’s net improvement, at 73% lower cost - 76% of Kimi K3 (Max)’s performance, at 92% lower cost. Against models as powerful as Fable 5 or stronger, DeepSeek-V4.1-Flash (Max) retains 35–54% of their net improvement at 97–99% lower cost. Those top models cost 37–76× more per task. Net improvement over Arena baseline | Median cost/task: - Claude Fable 5.1 (Max): +13.90% | $4.54 - GPT 6 Astra (Max): +11.90% | $4.09 - Claude Opus 5 (Max): +11.09% | $3.52 - Claude Opus 5 (High): +10.49% | $2.24 - Claude Fable 5 (High): +9.03% | $2.19 - Claude Opus 4.8 (High): +7.75% | $1.36 - GPT 5.6 Sol (xHigh): +7.40% | $1.09 - Kimi K3 (Max): +6.39% | $0.77 - Hy4 preview: +4.96% | $0.22 - DeepSeek-V4.1-Flash (Max): +4.87% | $0.06 With this release, GPT-5.6 Luna (xHigh), GLM-5.3-Flash, and DeepSeek-V4-Flash fell off the Pareto frontier for Agent Arena. Congrats again to the @deepseek_ai team on this release!
Show more
0
54
1.7K
145
Forward to community
Do models need their native harness for coding? Awesome work from our intern @melissapan on this. She dug into whether the harness (Claude Code vs Codex CLI vs Pi) actually moves the needle for coding agents. Turns out it matters way less than people assume. 21 model-harness pairs, real rigor. More from Arena to come. More details on the Arena blog:
Show more
Does your Claude model really need Claude Code…? 🤔 We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising findings emerge: 1️⃣Harness choice has little effect on task success rate, but can significantly affect the cost 2️⃣A simple harness can be competitive 3️⃣The native harness isn’t always the best. Millions of people are using coding agents, but the impact of harness choice remains unclear. (1/n) More details in the thread. 🧵
Show more
Arena Conversations: @petergostev sat down with @mattaitken, founder and CEO of @triggerdotdev, to discuss building infrastructure for long-running agents, including retries, API reliability, and why customers are switching frontier models faster than ever. Watch the full episode at the link below.
Show more
Check out the full Agent Arena leaderboard and Pareto frontier at:
DeepSeek-V4.1-Flash (Max) is a breakthrough in performance to cost efficiency. With +4.87% net improvement at $0.07 cost per median task, it’s reshaped the Pareto frontier for Agent Arena! Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest median task cost. For comparison, it retains: - 98% of Hy4 preview’s net improvement, at 73% lower cost - 76% of Kimi K3 (Max)’s performance, at 92% lower cost. Against models as powerful as Fable 5 or stronger, DeepSeek-V4.1-Flash (Max) retains 35–54% of their net improvement at 97–99% lower cost. Those top models cost 37–76× more per task. Net improvement over Arena baseline | Median cost/task: - Claude Fable 5.1 (Max): +13.90% | $4.54 - GPT 6 Astra (Max): +11.90% | $4.09 - Claude Opus 5 (Max): +11.09% | $3.52 - Claude Opus 5 (High): +10.49% | $2.24 - Claude Fable 5 (High): +9.03% | $2.19 - Claude Opus 4.8 (High): +7.75% | $1.36 - GPT 5.6 Sol (xHigh): +7.40% | $1.09 - Kimi K3 (Max): +6.39% | $0.77 - Hy4 preview: +4.96% | $0.22 - DeepSeek-V4.1-Flash (Max): +4.87% | $0.06 With this release, GPT-5.6 Luna (xHigh), GLM-5.3-Flash, and DeepSeek-V4-Flash fell off the Pareto frontier for Agent Arena. Congrats again to the @deepseek_ai team on this release!
Show more
0
54
1.7K
145
Forward to community
Coming to SF for Tech Week or COLM? We’d love to have you at Arena’s: Rooftop Happy Hour on 10/7. We're curating a group of researchers, developers, and builders pushing on hard AI problems to come take a break from the conference rooms up on our new office rooftop! If that's you, come hang out, meet the team, and check out our new office. Boba bar, light bites, refreshments and more will be provided. Space is limited. Register to request a spot on the list!
Show more
Exciting news: DeepSeek-V4.1-Flash (Max) by @deepseek_ai just landed in Agent Arena at #3# among open models! With +4.87% net improvement and a median cost per task of $0.07 it reshaped the Pareto frontier. Among the top 3 open models, DeepSeek-V4.1-Flash (Max) has the lowest median cost per task. Its +4.87% net improvement is within 0.09 percentage points of Hy4 preview (ranked #2#) at 68% lower cost, and within 1.52 percentage points of Kimi K3 (Max) (ranked #1#) at 91% lower cost. - Kimi K3 (Max): +6.39% | $0.77/task - Hy4 preview: +4.96% | $0.22/task - DeepSeek-V4.1-Flash (Max): +4.87% | $0.07/task See the full Pareto Frontier below. DeepSeek-V4.1-Flash (Max) is ranked #12# overall, and by signal landed #4# Confirmed Success with +13.75%! Congrats to the @deepseek_ai team on this release!
Show more
We analyzed how similar model responses were across 30,086 Arena battles. Models shared 43% of their ideas on average. We might expect that models from the same lab, or country, would show greater conceptual overlap. But the results don’t consistently support that. Claude Fable 5 illustrates this pattern: its closest conceptual match was neither Opus nor Sonnet. Which model came closest, along with the broader findings, may surprise you. Check out the full article from @DawidGalarowicz and @petergostev below.
Show more
When we started @arena as a project at UC Berkeley, we were ranking eight models. Today we evaluate the performance of over 500 models, with new releases arriving every week. It was great catching up with @mamoonha and @Joubinmir to talk Arena’s mission, and how it’s helping power frontier AI capabilities made for real-world use. Thanks for having me on Grit @kleinerperkins
Show more
Our CEO and Co-Founder, @ml_angelopoulos, joined @kleinerperkins for a new episode of Grit. They discuss @arena's real-world vantage point on the rapid pace of AI progress: including how open models are closing key performance gaps, Chinese labs are advancing the frontier, and more.
Show more
Agent Arena measures models on millions of real-world, long-horizon agentic tasks. Models get web search, filesystem, and terminal tools to complete complex workflows: writing code, creating slide decks, researching the web, building apps, and analyzing documents. We use causal tracing methodology to measure a model's net improvement, which indicates how much it improves outcomes relative to the average model. Dive into Agent Arena at: and check out the Pareto frontier:
Show more
Last month’s Hy4 preview launch landed @TencentHunyuan the #2# spot on the top 10 labs in Agent Arena among open-source labs! Across 14.5K+ real-world agent sessions, Hy4 preview ranks #2# among open models (#10# overall) with a +5.4% net improvement, trailing only Kimi K3 (Max) at +6.6%—a gap of just nearly 1 percentage point. This level of performance comes at roughly 68% lower cost: $0.26 per median task, compared with $0.80 for Kimi K3 (Max). Hy4 preview is now on the Agent Arena Pareto frontier. See the link below to explore its placement. By signal, Hy4 preview excels in: - Confirmed Success: 12% — explicit user feedback that the task worked - Praise vs. Complaint: 9% — implicit sentiment in user reactions - Bash Recovery: 7.5% — recovery from CLI errors It shows no issues with Tool Hallucination — calling tools that don’t exist By category among open models, it’s also #2# in Code and Work, and #3# in Chat. With this strong performance at a competitive price, Hy4 preview is a huge contribution to the open-source ecosystem. Congrats to the @TencentHunyuan team on this strong release!
Show more
🚀 Hy4 preview is here. 770B, 49B active, 1M context. Built for productivity. Open source frontier. Consistent affordable price. Use it. Tell us what breaks. More on Hy blog: HuggingFace: Github:
Show more
We analyzed how @claudeai's writing has changed from Fable 5 to Fable 5.1 across tens of thousands of high-reasoning Text Arena outputs. Overall, Fable 5.1 by @AnthropicAI uses fewer agreement openers, fewer em dashes and less wording like “honestly” and “frankly”, while its answers have grown longer. Compared with Fable 5: – Agreement openers such as “yes” and “exactly” are 58% less common, appearing in 0.99% of responses versus 2.35%. – Honesty wording such as “honestly” and “frankly” falls 45% per 1,000 words. – Em dashes fall 32% per 1,000 words. It has also shifted some patterns in another direction, e.g.: semicolons increase 63% per 1,000 words. Fable 5.1 still writes long answers, with sentence length and clause frequency close to Opus 5: – Median response length rises 30% from Fable 5: 319 → 414 words. – That remains 21% shorter than Opus 5’s 525 words. – Its sentence-length measure is slightly higher than Opus 5’s: 12.54 versus 12.20 words per sentence. Other familiar patterns also become less common compared with Fable 5: – Stock phrases such as “load-bearing” fall 20% per 1,000 words. – Hedges and caveats such as “perhaps” and “arguably” fall 36% per 1,000 words. – Praise and validation appear in 1.98% of responses, down from 3.17%. Long content words fall from 42.6% to 38.6%, while abstract nouns fall 25%, from 4.39 to 3.28 per 100 words. Let us know what you think in the comments.
Show more
We analyzed both @claudeai Fable 5 and Fable 5.1 on writing clarity ranked by humans, and are sharing results later today. Which model do you think writes more clearly?
Game development remains one of the most-requested, and most-challenging, categories on Arena. @iamwaynechi, PhD candidate at Carnegie Mellon University and research intern at Arena, just walked us through GameDevBench: a benchmark built from real tutorials that turns game development into verifiable, deterministic tasks. How do top frontier models perform on tasks a human beginner could complete in under an hour, and is the biggest bottleneck coding or multimodal understanding? Check out Wayne’s full talk to hear results and learn more about GameDevBench:
Show more