Register and share your invite link to earn from video plays and referrals.

Artificial Analysis
@ArtificialAnlys
Independent analysis of AI
682 Following    150.8K Followers
As with our specification for Terminal-Bench 4.0, we run Terminal-Bench-Science with the mini-swe-agent harness across all models. This keeps comparisons like-for-like, and in our testing often retains comparable evaluation performance to first-party harnesses like Codex and Claude Code. We're excited to keep updating the leaderboard as the frontier shifts. The Terminal-Bench-Science team’s call for contributions for 0.2 is open until October 5 at Explore the full results at
Show more
Recent releases from OpenAI and Anthropic lead other models by a wide margin on Terminal-Bench-Science. GPT-6 Astra and Claude Opus 5.5 have a ~20-point lead over Fable 5.1 in our testing. The best-performing model we’ve evaluated so far outside of these labs is Qwen3.8 Max (0902) at 12%. Within the same model families, the most recent GPT and Claude model releases made large gains. Comparing max effort, GPT-6 Sol and Opus 5.5 both saw performance improvements on the dataset while also reducing cost per task compared to GPT-5.6 Sol and Opus 5.
Show more
Terminal-Bench-Science discriminates well between models, as well as between different reasoning effort levels within frontier models. Claude Opus 5.5 rises ~38 points from low effort at 24% to xhigh at 62%, alongside a 5x difference in the cost per task. Max effort for Opus 5.5 scores slightly below xhigh at 59%. GPT-6 Sol gains 27 points from low to max effort at ~7.5x the cost per task.
Show more
Today we’re launching our leaderboard for Terminal-Bench-Science 0.1, an agentic benchmark for scientific research work. GPT-6 Astra (max) and Claude Opus 5.5 (xhigh) currently top it at 63% and 62% Terminal-Bench-Science 0.1 was announced in August 2026, built by @StevenDillmann and researchers at Stanford with the @terminalbench team and a global community of open source scientific contributors. It has 70 expert-curated tasks based on real research in five domains: life sciences (19), physical sciences (17), mathematical sciences (17), engineering sciences (9), and earth sciences (8). As with Terminal-Bench, each task drops an agent into a sandbox environment with the data, tools and instructions for a task, and the agent runs from start to finish. Automated tests grade every task pass/fail, and we report the average pass@1 over 3 attempts. Key takeaways: ➤ Only two models score above 50%: GPT-6 Astra (max) at 63%, and Claude Opus 5.5 at 62% (xhigh) and 59% (max). Recent releases have made large advancements, but even the top model, GPT-6 Astra, has significant headroom ➤ Life sciences is the lowest-scoring domain for most leading models, though domain scores are noisy with 8 to 19 tasks each. Claude Opus 5.5 (xhigh) passes 71% of mathematical sciences tasks but 46% of life sciences tasks ➤ The best open weights models, GLM-5.3 (max) at 10% and DeepSeek V4.1 Flash (max) at 9%, sit more than 50 points below the leaders
Show more
The Intelligence Index vs Cost per Task Pareto frontier shifted this week with the releases of MiMo-V2.6-Pro, Claude Opus 5.5, GPT-6 Luna, and GPT-6 Sol Together they have established eleven new points on the Pareto frontier (driven by different reasoning efforts): five from GPT-6 Luna, one each from MiMo-V2.6-Pro and GPT-6 Sol, and four from Claude Opus 5.5. GPT-6 Luna (max) scores 37 at $0.068 per task, MiMo-V2.6-Pro scores 46 at $0.13, GPT-6 Sol (max) scores 48 at $1.06, and Claude Opus 5.5 (max with fallback) is the new highest-scoring model at 58 at $5.98.
Show more
Explore the multilingual Controlled Voice Arena Leaderboards and compare the leading Text to Speech models across all 9 new languages: Controlled Voice Arena Leaderboard - select a language: Vote in the public Text to Speech Arena: Read our Text to Speech benchmarking methodology:
Show more
Open Weights No single Open Weights model dominates across languages, with four different labs leading depending on the language: ➤ Boson AI's Higgs Audio V3 TTS is the top Open Weights model across 4 of the 9 new languages: Hindi, German, Portuguese and Vietnamese. ➤ @FishAudio OpenAudio S1 Mini leads Open Weights in Japanese, Spanish and Arabic. ➤ @MistralAI Voxtral TTS leads Open Weights in French. ➤ BreezeBlue Breeze TTS 2 leads Open Weights in Mandarin. Open Weights coverage remains much thinner than proprietary models, with just 2 to 6 Open Weights models ranked on each language leaderboard, compared with 15 to 24 public models overall.
Show more
European Languages 🇪🇸 🇩🇪 🇫🇷 🇵🇹 ➤ Cartesia ranks #1# across all four European language leaderboards, with Sonic 3.5 leading Portuguese and Sonic 3.6 leading Spanish, German and French. ➤ Inworld's Realtime TTS-2 ranks #2# in German at 1,261 Elo, just 22 points behind Sonic 3.6. It also ranks #4# in French and Portuguese. ➤ Portuguese is the only new language where Sonic 3.5 leads Sonic 3.6, with a 33-point Elo difference between the two models. Spanish: Sonic 3.6 ranks #1# at 1,229 Elo, followed by Eleven v3 at 1,202 and Sonic 3.5 at 1,161. Realtime TTS-2 ranks #4# at 1,145, followed by v3 Conversational at 1,134. German: Sonic 3.6 ranks #1# at 1,283 Elo, followed by Realtime TTS-2 at 1,261 and Sonic 3.5 at 1,221. Eleven v3 and v3 Conversational round out the top 5 at 1,198 and 1,191. French: Sonic 3.6 ranks #1# at 1,349 Elo, leading by 71 points over v3 Conversational at 1,278. Eleven v3 ranks #3# at 1,259, followed by Realtime TTS-2 at 1,258 and Sonic 3.5 at 1,238. Portuguese: Sonic 3.5 ranks #1# at 1,324 Elo, ahead of Sonic 3.6 at 1,291 and Eleven v3 at 1,283. Realtime TTS-2 ranks #4# at 1,278, followed by v3 Conversational at 1,262.
Show more
Hindi & Arabic 🇮🇳 🇸🇦 ➤ Cartesia's Sonic 3.6 leads both Hindi and Arabic, with a 33-point Elo lead in Hindi and a 132-point lead in Arabic over the next-ranked model. ➤ ElevenLabs' Eleven v3 ranks #2# in both languages, with v3 Conversational ranking #3# in both. Hindi: Cartesia's Sonic 3.6 ranks #1# at 1,179 Elo, followed by ElevenLabs' Eleven v3 at 1,146 and v3 Conversational at 1,126. Sonic 3.5 ranks #4# at 1,117, followed by Inworld's Realtime TTS-2 at 1,066. Arabic: Sonic 3.6 ranks #1# at 1,369 Elo, followed by Eleven v3 at 1,237 and v3 Conversational at 1,223. SpaceXAI TTS ranks #4# at 1,209, followed by ElevenLabs' Multilingual v2 at 1,184.
Show more
East & Southeast Asia 🇯🇵 🇨🇳 🇻🇳 ➤ Cartesia's Sonic 3.6 leads Japanese and Vietnamese, while Inworld's Realtime TTS-2 ranks #1# in Mandarin. ➤ StepFun's StepAudio 2.5 TTS ranks #3# in Mandarin at 1,130 Elo, just 16 points behind Sonic 3.6. ➤ In Japanese, Inworld's Realtime TTS-2 ranks #2# at 1,123 Elo, ahead of ElevenLabs' v3 Conversational and Sonic 3.5. Japanese: Cartesia's Sonic 3.6 ranks #1# at 1,153 Elo, followed by Inworld's Realtime TTS-2 at 1,123 and ElevenLabs' v3 Conversational at 1,095. Sonic 3.5 and Eleven v3 round out the top 5 at 1,093 and 1,085. Mandarin: Inworld's Realtime TTS-2 ranks #1# at 1,185 Elo, followed by Sonic 3.6 at 1,146 and StepFun's StepAudio 2.5 TTS at 1,130. Alibaba's Qwen-Audio-3.0-TTS-Plus and MiniMax's Speech 2.8 Turbo round out the top 5. Vietnamese: Sonic 3.6 ranks #1# at 1,721 Elo, followed by ElevenLabs' v3 Conversational at 1,677 and Eleven v3 at 1,676. Sonic 3.5 ranks #4# at 1,650, followed by Realtime TTS-2 at 1,600.
Show more
Announcing Multilingual Text to Speech Arena Leaderboards, comparing leading TTS models across 9 languages beyond English 🇯🇵 🇨🇳 🇮🇳 🇪🇸 🇩🇪 🇫🇷 🇵🇹 🇻🇳 🇸🇦 Text to Speech performance varies significantly across languages, with models that perform well in English not necessarily delivering the same pronunciation, pacing, tone, and naturalness in other languages. We have extended the Artificial Analysis Controlled Voice Arena to 9 new languages: Japanese 🇯🇵, Mandarin Chinese 🇨🇳, Hindi 🇮🇳, Spanish 🇪🇸, German 🇩🇪, French 🇫🇷, Portuguese 🇵🇹, Vietnamese 🇻🇳, and Arabic 🇸🇦. Each model is evaluated using standardized cloned voices, with prompts written natively in each language and preference votes collected from first-language speakers. Elo scores are calculated independently for each language, allowing each leaderboard to reflect model preferences among speakers of that language. Key results: ➤ @Cartesia's Sonic family leads 8 of the 9 new language leaderboards, with Sonic 3.6 ranking #1# in 7 languages and Sonic 3.5 leading Portuguese. @InworldAI's Realtime TTS-2 takes the #1# spot in Mandarin. ➤ @ElevenLabs' Eleven v3 family ranks in the top 3 across 7 of the 9 new languages, through Eleven v3 and Eleven v3 Conversational. ➤ Inworld's Realtime TTS-2 leads Mandarin at 1,185 Elo, ahead of Sonic 3.6 at 1,146 and StepFun's StepAudio 2.5 TTS at 1,130. Realtime TTS-2 also ranks #1# in English (US). ➤ More than 100,000 human preference votes have been collected across the 9 new language leaderboards, with 15 to 24 public models ranked per language. See the per-language results below ⬇️
Show more
Compare GPT-6 Sol and Luna with other leading models at:
Breakdown of the individual evaluations in the Artificial Analysis Intelligence Index v4.3
Both models regress in GDPval-AA v2.1 at max effort, driven by shorter deliverables that more often omit required elements.
GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and regressions in others Pricing is approximately half that of GPT-5.6: Sol drops from $4/$20 to $2/$10 per million input/output tokens, and Luna from $0.20/$1.20 to $0.10/$0.50, with the same 90% discount for cache reads and 25% premium for cache writes. Key takeaways: ➤ Halves Cost per Task: GPT-6 Sol (max) costs $1.06 per task to run the Artificial Analysis Intelligence Index, ~50% less than GPT-5.6 Sol (max) at $1.99. GPT-6 Luna (max) costs $0.07 per task, ~60% less than GPT-5.6 Luna (max) at $0.18. This is driven by the price cut, as both models use slightly more output tokens per task (31k vs 29k for Sol, and 51k vs 41k for Luna). These two releases allow OpenAI to capture a significant portion of the cost efficiency Pareto frontier. ➤ In the Coding Agent Index, Sol improves but Luna regresses: In OpenAI's Codex harness, GPT-6 Sol (max) scores 57 in the Artificial Analysis Coding Agent Index, up 2 points from GPT-5.6 Sol (max), with gains in Terminal-Bench 4.0 (43% vs 37%) and SWE-Atlas-QnA (58% vs 54%). At $2.99 per task it costs ~50% less than GPT-5.6 Sol (max) and sits on the Pareto frontier of Coding Agent Index vs Cost per Task. GPT-6 Luna (max) scores 41, down 2 points from GPT-5.6 Luna (max), with lower scores in SWE-Atlas-QnA (44% vs 49%) and DeepSWE v1.1 (64% vs 66%), at ~60% lower cost per task. ➤ Significant reduction in hallucination: Both models hallucinate less in AA-Omniscience, our knowledge and hallucination benchmark. GPT-6 Sol (max) cuts its hallucination rate from 92% to 60% and GPT-6 Luna (max) from 93% to 77%. Sol achieves this by declining to answer more often: it attempts 83% of questions vs 99% for GPT-5.6 Sol (max), which cuts wrong answers by about a quarter but also lowers accuracy 5 points from 59% to 54%. Luna's accuracy is broadly unchanged at 44% vs 43% while it answers fewer questions. On the AA-Omniscience Index, Sol improves from 22 to 27 and Luna from -10 to 1. ➤ Mix of improvement and regression across evals: Beyond AA-Omniscience, both models improve in AutomationBench-AA (Sol 62% vs 60%, Luna 53% vs 50%) and Terminal-Bench 4.0 (Sol 44% vs 40%, Luna 13% vs 12%). However, we observe regressions in two key knowledge work evaluations. In GDPval-AA v2.1, our benchmark adapted from OpenAI's dataset of economically valuable tasks across 44 occupations, Sol drops ~100 Elo points and Luna ~75. Luna also drops ~45 Elo points in AA-Briefcase v1.1, while Sol is level. AA-Briefcase v1.1 is a private evaluation across multi-week knowledge work projects, with thousands of input files. Our team has manually inspected hundreds of model outputs: the regressions tend to be driven by reduced presentation quality and deliverables that omit rubric elements. Congratulations @OpenAI and @sama on the launch!
Show more
0
116
1.7K
148
Forward to community
Compare Claude Opus 5.5 with other leading models at:
Full breakdown of the individual evaluations in the Artificial Analysis Intelligence Index for Claude Opus 5.5 across all reasoning efforts
Claude Opus 5.5 (max) is a top performer in agentic knowledge work. It leads AA-Briefcase v1.1 at 1,822 Elo (+143 over Fable 5.1) and GDPval-AA v2.1 at 1,846 Elo (+111 over Claude Fable 5.1, +138 over Claude Opus 5). On AA-Briefcase it leads across analytical quality and presentation sub-scores, and sits just behind Fable 5.1 for rubric-based scoring
Show more
Claude Opus 5.5 takes the top spot on the Artificial Analysis Intelligence Index, along with a 20% price cut and larger cache hit discount Claude Opus 5.5 brings Anthropic to parity with GPT-6 Astra on evaluations like Terminal-Bench 4.0 and AutomationBench-AA, while extending Anthropic’s lead in agentic knowledge work. At max effort it scores 58 on the Artificial Analysis Intelligence Index, the highest score we have measured by several points. Anthropic has cut Opus pricing to $4/$20 per 1M input/output tokens (Opus 5: $5/$25) and cache reads from $0.50 to $0.20. Key takeaways: ➤ Consistent strong performance, with leading scores on six of the ten Intelligence Index evaluations: Humanity's Last Exam 61.4% (previous best 59.1%, Claude Fable 5.1), SciCode 66.9% (63.1%, Fable 5.1), GDPval-AA v2.1, AA-Briefcase v1.1, AA-Omniscience and AutomationBench-AA. On Terminal-Bench 4.0 it scores 59.6%, level with the leader GPT-6 Astra (xhigh) and +11 points over Opus 5. It remains slightly behind on CritPt, AA-LCR, and GDP.pdf ➤ Leads in agentic knowledge work: On AA-Briefcase, our private frontier knowledge work evaluation, it reaches an Elo of 1822. This is +143 over Fable 5.1, ahead on both analytical quality and presentation, and is the first time Anthropic has reached presentation quality surpassing GPT-5.6 Sol. This evaluation tests whether models can produce accurate and well-presented professional outputs using our open source reference agent harness, Stirrup ➤ Level with Opus 5 on cost per task despite 1.6x the output tokens: Opus 5.5 (max) uses ~119k output tokens per Intelligence Index task, against ~73k for Opus 5 (max), ~78k for Fable 5.1 (max) and ~27k for GPT-6 Astra (max) ➤ Four of five effort levels sit on the Intelligence vs Cost per Task frontier: Opus 5.5 max, xhigh, high, and medium all sit on the Pareto frontier, costing less or outperforming other models scoring 50+ (GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5) Other model details: ➤ Context window: 1 million token context with image and text input support, unchanged from Opus 5 ➤ Pricing: $4/$20 per 1M input/output tokens, down 20% from $5/$25 for Opus 5. Cache writes $5 per 1M tokens for the 5 minute TTL, down from $6.25. Cache reads have been further discounted to $0.20 per 1M tokens, down 60% from Opus 5’s $0.50. This is a 95% discount compared to uncached input pricing, up from 90% on previous Opus models ➤ Effort settings: Five effort settings (low, medium, high, xhigh, and max). Intelligence Index evaluations were run at all five with Anthropic's default fallback enabled
Show more
0
121
3K
280
Forward to community
StepFun's Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index, matching Kimi K3 (max) at ~2.8x lower cost per task, but trails peers on agentic evaluations Step 5 Preview is @StepFun_ai's new flagship model, with 600B total and 27B active parameters, succeeding Step 3.7 Flash (released May 2026). It scores 44 on the Intelligence Index, level with Kimi K3 (max) and just behind GLM-5.3 (max, 45) and Qwen3.8 Max (45) Key takeaways: ➤ Step 5 Preview costs ~2.8x less per Intelligence Index task than models at the same score. It costs ~$0.72 per task, against ~$2.00 for Kimi K3 (max) at the same score of 44 and ~$2.01 for GLM-5.3 (max) at 45. This is driven by pricing: at $1/$2.70 per 1M input/output tokens, it is priced below both on input and output. MiMo-V2.6-Pro is the one model that scores higher (46) at a lower cost per task ($0.13) ➤ Frontier reasoning is the standout strength, and where the jump from Step 3.7 Flash is largest. Step 5 Preview scores 46% on Humanity's Last Exam, in line with Kimi K3 (max, 47%), and 21% on CritPt, between Kimi K3 (23%) and GLM-5.3 (max, 19%). Both are up sharply from Step 3.7 Flash: +25 points on HLE and +19 points on CritPt ➤ Higher AA-Omniscience accuracy than GLM-5.3 at fewer parameters, but with more hallucination. At 600B total parameters, Step 5 Preview reaches 42% accuracy on AA-Omniscience, our benchmark measuring factual recall and hallucination, ahead of GLM-5.3 (max, 34%, 753B) and behind Kimi K3 (max, 48%, 2.8T). It attempts more questions than GLM-5.3 (68% vs 55%) and hallucinates more often when it does (43% vs 30%), landing at 16 on the AA-Omniscience Index, between GLM-5.3 (14) and Kimi K3 (20) ➤ Agentic evaluations are where Step 5 Preview lags peers at a similar Intelligence Index score. It scores 1,566 Elo on GDPval-AA, our primary evaluation for agentic performance, behind Qwen3.8 Max (1,668) and GLM-5.3 (max, 1,646). The gap holds on Terminal-Bench 4.0 (33% vs 39% and 42%), AA-Briefcase (1,432 Elo vs 1,640 and 1,525) and AutomationBench-AA (51% vs 56% and 62%) Key model details: ➤ Model Size: 600B total parameters, 27B active MoE model ➤ Context window: 1M tokens ➤ Multimodality: Text, image and video input, text output ➤ Pricing: $1/$2.70 per 1M input/output tokens, with cached input at $0.05/M ➤ Availability: StepFun first-party API, with open weights release planned for October 15th ➤ Licensing: Closed weights currently, with weights release planned for October 15th
Show more