Register and share your invite link to earn from video plays and referrals.

Artificial Analysis
@ArtificialAnlys
Independent analysis of AI
664 Following    125.2K Followers
Ant Group has just released Ling 3.0 Flash, a 124B open weights model that scores 38 on the Artificial Analysis Intelligence Index. Ling 3.0 demonstrates a marked improvement over the previous generation and sits on the Pareto frontier for Intelligence versus Total Parameters among open weights models @AntGroup has released Ling 3.0 Flash, an open weights reasoning model with 124B total parameters and 5B active at inference time and a 262K token context window. It scores 38 on the Artificial Analysis Intelligence Index v4.1.1, 24 points above the previous generation Ling 2.6 Flash (Non-reasoning, 14). This matches MiMo-V2.5 (38) and Qwen3.6 27B (38) while using a third of MiMo-V2.5's active parameters. It remains behind the flash-tier open weights leader, DeepSeek V4 Flash 0731 (Reasoning, Max Effort) at 52. Key results: ➤ Ling 3.0 Flash sits on the open weights Pareto frontier for Intelligence vs. Total Parameters. No open weights model with fewer than 124B total parameters scores higher on the Artificial Analysis Intelligence Index, and at a comparable total size gpt-oss-120b (117B) scores 24, 14 points behind. The next model up the frontier is MiniMax-M2.7, which scores 39 with 230B total parameters. ➤ Ling 3.0 Flash demonstrates meaningful improvements in agentic abilities. Ling 3.0 Flash scores 27% on τ3-Bench Banking, second among flash-tier open weights models behind DeepSeek V4 Flash 0731 (Reasoning, Max Effort) at 39%, and ahead of Hy3 (23%) and Inkling Small (19%), all of which score 3 to 4 points higher on the Index. On GDPval-AA v2, it reaches an Elo rating of 1108, which is a meaningful improvement over Ling 2.6 Flash (545) ➤ Ling 3.0 Flash makes marked improvement in Aa-Omniscience, but almost entirely through abstention rather than knowledge. Ling 3.0 Flash scores -18 on the Artificial Analysis Omniscience Index, up from -66 for Ling 2.6 Flash (Non-reasoning). The underlying accuracy moved only from 16% to 18% while the hallucination rate on wrong answers fell from 97% to 44% and the attempt rate fell from 99% to 56%. Ling 3.0 Flash answers far fewer questions but is far less likely to hallucinate when it does answer ➤ Ling 3.0 Flash is on the Pareto frontier for Intelligence versus Price among similar sized open weights models. At $0.075 per 1M input and $0.22 per 1M output tokens on inclusionAI's first-party API, Ling 3.0 Flash is the cheapest model per token that we have measured at 38 or above on the Intelligence Index. However, that advantage is dampened in Cost per Task, because Ling 3.0 Flash used ~240M output tokens to run the Intelligence Index, at a total cost of $73. At $0.02 per task it sits inside the Pareto frontier for intelligence versus Cost per Task rather than on it. Additional model details: ➤ Size: 124B total parameters, 5B active ➤ Context window: 262K ➤ Pricing: $0.075 per 1M input tokens and $0.22 per 1M output tokens with an 80% cache hit discount ➤ License: MIT ➤ Providers: inclusionAI first-party API and DeepInfra third-party API
Show more
Muse Spark 1.2 places Meta on the Cost per Task Pareto frontier, scoring 6 points below Claude Opus 5 at ~1/6th of the cost At Meta's $1.25/$4.25 per 1M token pricing, Muse Spark 1.2 (xhigh) sits on the Pareto frontier of Intelligence Index vs Cost per Task. It delivers comparable intelligence to Claude Opus 4.8 (max, $2.03) at a fifth of the cost per task, and undercuts GPT-5.6 Sol (high, $0.55), GPT-5.6 Terra (max, $0.61), and Kimi K3 (max, $0.87). The nearest cheaper options are Grok 4.5 (high, $0.36) and GPT-5.6 Sol (medium, $0.37), and they all sit below it on the Index. The step up from Muse Spark 1.1 ($0.29 per task) comes at unchanged per-token pricing, with the increase driven by heavier token usage on agentic tasks.
Show more
Alibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index at $1.14 per task, but open weights leader Kimi K3 remains 1 point ahead at 25% lower cost per task ($0.86) @Alibaba_Qwen has released Qwen3.8 Max, which Alibaba states is a 2.4T total parameter MoE activating 95B parameters per forward pass. Alibaba has announced it plans to release the weights next week, a shift in strategy as it has typically kept its Max class of models proprietary. Once released, Qwen3.8 Max would be ~6x larger than Alibaba's largest open weights release to date (Qwen3.5 397B) and the second largest open weights model behind Kimi K3 (2.8T) Note: we earlier published results showing Qwen3.8 Max scoring 53 on the Artificial Analysis Intelligence Index. Those runs were affected by intermittent issues on the endpoint we were evaluating, and we have re-run all evaluations on Alibaba's public API endpoint Key takeaways: ➤ Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, up 10 points from Qwen3.7 Max (46). It is in line with Claude Opus 4.8 (max, 56) and sits second among labs from China, ahead of GLM-5.2 (max, 51) but behind Kimi K3 (max, 57) ➤ Qwen3.8 Max scores 1739 Elo on GDPval-AA, a 468 Elo gain over Qwen3.7 Max. This places it ahead of Kimi K3 (1685), effectively tied with Claude Fable 5 (1743) and GPT-5.6 Sol (max, 1730), and behind only Claude Opus 5 (max, 1852) ➤ Gains over Qwen3.7 Max span agentic evaluations, scientific reasoning and coding: Terminal-Bench v2.1 +6 points, CritPt +7 points, SciCode +4 points and HLE +3 points, with GPQA unchanged. AA-LCR (-2 points) and AA-Omniscience (-10 points, driven by a hallucination rate rising 23% to 40%) regress ➤ Qwen3.8 Max costs $1.14 per Intelligence Index task, more than double Qwen3.7 Max ($0.53), at ~1.3x Kimi K3 (max, $0.86) and ~2x GLM-5.2 (max, $0.57). Cost is driven in part by more turns on agentic evaluations, with GDPval-AA input tokens rising ~15x over Qwen3.7 Max, and output token usage up 45% to 145M ➤ The 𝜏³-Bench Banking result (42%) appears out of distribution. It is a 32 point gain over Qwen3.7 Max and places Qwen3.8 Max ahead of models that outscore it on other evaluations Key model details: ➤ Size: 2.4T total parameters, ~95B active per forward pass (MoE) ➤ Context window: 1M tokens ➤ Multimodal: text, image and video input with text output ➤ Pricing: $2.00/$6.00 per 1M input/output tokens on the @alibaba_cloud first-party API, with a $0.25 cache hit price. This is lower than Qwen3.7 Max across the board ($2.50/$7.50, with a $0.50 cache hit price) ➤ Availability: Alibaba Cloud first-party API. Alibaba states the weights will be released next week
Show more
Meta has released Muse Spark 1.2. It's their third release in four months and scores 54 on the Artificial Analysis Intelligence Index, significantly improving agentic knowledge work capabilities over prior releases and putting Meta next to SpaceXAI in a tie for third place amongst US labs Muse Spark 1.2 (xhigh) lands at 54, up 3 points from Muse Spark 1.1 (51) and 11 points from Muse Spark 1.0 (43, April). It enters effectively tied with GPT-5.5 (xhigh, 55) and Grok 4.5 (high, 54), narrowly behind current frontier models Claude Opus 5 (max, 61), Claude Fable 5 (max w/ fallback, 60), GPT-5.6 Sol (max, 59), and Kimi K3 (max, 57) Congratulations to @AIatMeta, @finkd, and @alexandr_wang on the release! Key Takeaways: ➤ Muse Spark 1.2 gets closer to the frontier on agentic knowledge work. At Muse Spark 1.1's launch, we noted agentic knowledge work as its clearest gap; Muse Spark 1.2's gains help to close this. Its GDPval-AA v2 Elo rose 260 points to 1631, #5# among all models we have benchmarked and ahead of Claude Opus 4.8 (max, 1588). Terminal-Bench 2.1 gained 2 points (78% to 80%), and Tau3-Bench Banking rose 2 points (25% to 27%) ➤ Among the most cost-efficient models at its intelligence level. Muse Spark 1.2 costs $0.40 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing, with only Grok 4.5 (high, $0.37) and GPT-5.6 Sol (medium, $0.39) cheaper in its intelligence cluster - GPT-5.6 Terra (max, $0.51), Kimi K3 (max, $0.86), and GPT-5.5 (xhigh, $1.18) all cost more per task. The cost increase over Muse Spark 1.1 ($0.29 per task) is driven by increased token usage per Intelligence Index task ➤ AA-Omniscience abstention rate increases. The score rose from 18 to 22 as the hallucination rate fell 10 points (38% to 28%) and the attempt rate dropped from 82% to 67%. This heavy abstention (not answering questions when unsure) now drives both the low hallucination rate and a lower accuracy (41% to 38%) ➤ Scientific Reasoning results remain largely unchanged. CritPt notably gained 3 points (15% to 18%), while SciCode fell 2 points (58% to 56%), and Humanity's Last Exam fell 1 point (45% to 44%) Other model details: ➤ Context window: 1M tokens, unchanged from Muse Spark 1.1 ➤ Pricing: unchanged from Muse Spark 1.1: $1.25/$4.25 per 1M input/output tokens, with cache hits discounted to $0.15 per 1M ➤ Availability: Meta's first-party API at launch
Show more
0
62
1.1K
107
Forward to community
Announcing the Artificial Analysis Endpoint Accuracy Index, measuring how much of an open weights model's accuracy each serverless API endpoint preserves. We are initiating coverage with GLM-5.2, gpt-oss-120b and DeepSeek V4 Pro, with Kimi K3 coming soon Providers trade off accuracy to optimize for speed and cost. They quantize weights, write custom kernels and tune their inference stacks, and sometimes they simply ship bugs. We are bringing the rigor of our Artificial Analysis Intelligence Index to measuring endpoints, so developers can pick providers on accuracy, not just price and speed We benchmark each serverless endpoint against our own self-hosted reference deployment of the official weights, where 100% represents matching the reference. An endpoint is at reference parity when its result falls within the 95% confidence interval of the reference. Coverage is live for GLM-5.2, gpt-oss-120b and DeepSeek V4 Pro, with Kimi K3 accuracy coverage launching soon Key elements of the Endpoint Accuracy Index: ➤ Three areas, equally weighted: tool calling (BFCL-500, 500 questions, 3 repeats), scientific reasoning (HLE-250, 250 questions, 10 repeats) and long context recall (AA-LCR-25, 25 questions, 10 repeats). Each subset separates endpoints on the serving choices that drive accuracy differences, with repeats sized for tight confidence intervals ➤ Reference deployment: we self-host the official weights at the lab's recommended precision, following the lab's serving recipe, and publish the complete commands for each reference ➤ Inference parameters: we run the model's highest supported reasoning mode and each endpoint's highest supported output length and context window ➤ Confidence intervals: the parity test accounts for uncertainty in both the endpoint's runs and the reference's runs ➤ Rotating coverage: models enter once sufficient number of providers serve them and exit when a newer version in the same family supersedes them. We benchmark new endpoints as providers launch them and refresh all listed endpoints periodically ➤ Point in time: each result carries the date it was measured, with multi-day benchmarks dated to their final day Key results for GLM-5.2 ➤ Output token limits restrict accuracy. Restrictive limits cut responses off before the model finishes reasoning, and the most restrictive endpoints score half the reference or less on HLE-250 Key results for gpt-oss-120b ➤ Tool call handling separates endpoints. Providers parse and format tool calls differently, and some endpoints score 22% on BFCL-500 against 37% for the reference ➤ Serving configuration changes what the model does at the same requested settings. Some endpoints produce far fewer reasoning tokens at the same configured level, and restricted context windows truncate long context tasks Key results for DeepSeek V4 Pro ➤ DeepSeek V4 Pro endpoints are more in line with the reference. Majority of the endpoints are at reference parity, and DeepSeek's own first-party endpoint scores slightly above the reference
Show more
DeepSeek V4 Flash 0731 is now open weights! @deepseek_ai has just released the weights for its new flash tier model, DeepSeek V4 Flash 0731. With a score of 50 on the Artificial Analysis Intelligence Index, it lands among the top 3 open weights models on the leaderboard. The weights are released under the MIT license, allowing unrestricted commercial use and modification. DeepSeek V4 Flash 0731 shares identical architecture and pricing with the earlier DeepSeek V4 Flash. At a size of 284B total parameters (13B active), released in mixed FP4/FP8 precision at ~167GB total file size, it lands on our Pareto frontier for Intelligence Index vs. Total Parameters. Among open weights models, DeepSeek V4 Flash 0731 delivers a significant leap in intelligence for its size class. DeepSeek V4 Flash 0731 is also available now through DeepSeek's first-party API. Check out Artificial Analysis to compare DeepSeek V4 Flash 0731 with other leading open weights and proprietary models:
Show more
DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index, a 10-point jump over DeepSeek V4 Flash (released April 2026) that puts it 6 points ahead of DeepSeek V4 Pro. It shares identical architecture and pricing with the earlier DeepSeek V4 Flash, and lands on our Pareto frontier for Intelligence vs Cost per Task @deepseek_ai’s DeepSeek V4 Flash 0731 is one Intelligence Index point behind GPT-5.6 Luna (max, 51). Even after OpenAI’s 80% price cut on GPT-5.6 Luna today, DeepSeek V4 Flash 0731’s Cost per Task on DeepSeek’s first-party API comes in at ~60% lower than GPT-5.6 Luna (max), a model with comparable intelligence. A key driver of this is DeepSeek’s ~98% cache hit discount on its first-party API, a significantly more aggressive discount than the 90% cache hit discount offered by most of the industry The new model is a significant step up from the previous generation, DeepSeek V4 Flash (40), and places the model within 1 point of GLM-5.2 (max, 51). It remains 7 points behind the open weights frontier set by Kimi K3 (max, 57). For additional context, this places the model in line with recently released Gemini 3.6 Flash (50) and 1 point behind Muse Spark 1.1 (xhigh, 51). DeepSeek is expected to release the model’s full weights in the coming weeks DeepSeek V4 Flash 0731 retains a 1M token context window, and its size remains unchanged from DeepSeek V4 Flash at 284B total parameters and 13B active at inference time Key results: ➤ Improvements in agentic performance: DeepSeek V4 Flash 0731 achieves an Elo rating of 1559 on GDPval-AA v2, our evaluation focused on agentic real-world work tasks, up from 1189 for the previous DeepSeek V4 Flash. Once weights are released this will be the second highest open weights score, behind Kimi K3 (max, 1687) and ahead of GLM-5.2 (max, 1510). Terminal-Bench 2.1 rises 17 points to 79% and τ³-Bench Banking 8 points to 31% ➤ Token usage falls 12% against the predecessor: DeepSeek V4 Flash 0731 used ~206M output tokens to run the Intelligence Index, against ~234M for the previous DeepSeek V4 Flash. The new variant is more token efficient, achieving a higher Intelligence Index with a lower number of total output tokens ➤ DeepSeek V4 Flash 0731 improves over its predecessor on every evaluation in the Intelligence Index: Alongside the agentic gains, CritPt gains 9 points to 17%, SciCode 5 points to 50%, Humanity's Last Exam 5 points to 37%, AA-LCR 3 points to 66% and GPQA Diamond 1 point to 91% ➤ AA-Omniscience improvements are driven by fewer hallucinations, rather than higher accuracy: DeepSeek V4 Flash 0731 achieves an AA-Omniscience Index of -16, a +7 improvement from its predecessor. This improvement is purely driven by a reduced hallucination rate, with overall accuracy (percentage correct) unchanged. Its AA-Omniscience Hallucination Rate is 84%, a 12 point decrease from its predecessor, and comparable to models such as GPT-5.6 Terra (max, 85%) and Mistral Medium 3.5 (82%) Additional model details: ➤ Context window: 1M tokens (equivalent to DeepSeek V4 Flash) ➤ Size: 284B total parameters (13B active) ➤ Input modalities: Text input and output only ➤ Accessibility: Available through DeepSeek’s first-party API ➤ Pricing: $0.14/$0.28 per 1M input/output tokens, unchanged from DeepSeek V4 Flash. Cache hit price of $0.0028 per 1M tokens, a 98% discount
Show more
0
110
2.5K
230
Forward to community
Temporary error with cache hit rate calculation - it was rectified a couple of minutes after your screenshot! 0731 is only marginally higher Cost per Task than the earlier version, and via the DeepSeek API with ~99% cache hit discount it is most certainly on our Pareto frontier for Intelligence vs. Cost per Task.
Show more
We are committed to pushing the model frontier across cost efficiency, capability, and speed. Starting today, we are reducing prices for GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20% , and offering a faster option for GPT-5.6 Sol in the API. Luna and Terra’s lower prices are reflected in how usage is counted in Codex and ChatGPT Work, so your usage goes further.
Show more
0
1.3K
18.8K
1.9K
Forward to community
AA-Doomscroll: An endless, randomly shuffled feed of charts from across Artificial Analysis. Have fun. Live on the site now:
SpaceXAI has released Grok Voice Think Fast 2.0 today, with the High reasoning variant debuting at #2# on the Artificial Analysis Speech to Speech Index at 82.9%, and #1# on Tau Voice for Agentic Performance at 56.5% - among the fastest models at 0.70s Time to First Audio Grok Voice Think Fast 2.0 is SpaceXAI's successor to Grok Voice Think Fast 1.0 (75.7% on the Speech to Speech Index). It is the only model in the Index's top five with an average Time to First Audio under 1 second, achieving 0.70 seconds vs. 1.14 seconds for the next fastest, GPT-Realtime-2 High. Key takeaways: ➤ Speech to Speech Index: Grok Voice Think Fast 2.0 High debuts at #2# at 82.9%, only behind Qwen Audio 3.0 Realtime Plus (84.1%) and ahead of GPT-Realtime-2.1 High (79.1%) and GPT-Realtime-2 High (77.2%). This is up 7.3 percentage points from Grok Voice Think Fast 1.0 (75.7%) ➤ Speech to Speech Index by Benchmark: On Tau Voice, Grok Voice Think Fast 2.0 High is the new leader at 56.5%, just ahead of Qwen Audio 3.0 Realtime Plus at 54.6% and its predecessor Grok Voice Think Fast 1.0 at 52.1%. On Big Bench Audio, it achieves 97.2%, behind leader Qwen Audio 3.0 Realtime Plus at 99.2%. On our Full Duplex Bench subset, it scores 95.1%, up from 77.8% for Grok Voice Think Fast 1.0, the largest driver of its index gain, behind Qwen Audio 3.0 Realtime Plus at 98.4% ➤ Speed: The model’s average Time to First Audio on Big Bench Audio is 0.70 seconds, faster than GPT-Realtime-2 High (1.14s), GPT-Realtime-2.1 High (1.21s), and Grok Voice Think Fast 1.0 (1.25s), and well ahead of Qwen Audio 3.0 Realtime Plus (4.02s) ➤ Price: Grok Voice Think Fast 2.0 is priced at $4.80 per hour of input audio, up from $3.00 for Grok Voice Think Fast 1.0 and more expensive than Qwen Audio 3.0 Realtime Plus ($4.42) and GPT-Realtime-2 High ($4.14), but ~2.2x cheaper than GPT-Realtime-2.1 High ($10.75) Congratulations @SpaceXAI @elonmusk! See below for more detail ⬇️
Show more
OpenAI has released GPT Transcribe: a Speech to Text model scoring 3.31% on AA-WER (#9#), improving 0.7 p.p. over its predecessor GPT-4o Transcribe while lowering price 25% to $4.50 per 1,000 minutes of audio GPT Transcribe is OpenAI's latest non-streaming (batch) speech transcription model, now accepting three kinds of context to improve transcription quality: a text prompt describing the recording's topic or setting, keywords for literal terms that may appear in the audio (such as product names or acronyms), and multiple language hints for multilingual and code-switching audio. The model processes audio at ~34× real-time and is available at $4.50 per 1,000 minutes of audio ($0.0045/min) via the OpenAI API Platform. OpenAI has also released GPT-Live-Transcribe, a streaming Speech to Text model. We are currently benchmarking this model and plan to share results on our Streaming Speech to Text leaderboard. See more details below ⬇️
Show more
Alibaba has released Qwen Audio 3.0 Realtime, with the Plus variant debuting as the new #1# model on the Artificial Analysis Speech to Speech Index at 84.1%, ahead of GPT-Realtime-2.1 High at 79.1% Released earlier this month, Qwen Audio 3.0 Realtime is @Alibaba_Qwen’s flagship native Speech to Speech model, available in two variants: Plus and Flash. Qwen Audio 3.0 Realtime Plus leads on all three component benchmarks comprising the Artificial Analysis Speech to Speech Index, Big Bench Audio for Speech Reasoning, Full Duplex Bench for Conversational Dynamics, and Tau Voice for Agentic Performance. We tested the China-hosted endpoints on Aliyun (Alibaba Cloud). Key takeaways: ➤ Speech to Speech Index: Qwen Audio 3.0 Realtime Plus is the new leader at 84.1%, ahead of GPT-Realtime-2.1 High (79.1%) and GPT-Realtime-2 High (77.2%). The Flash variant comes in at 4th at 76.3%. ➤ Speech to Speech Index by Benchmark: On Big Bench Audio, the Plus variant achieves 99.2%, up ~0.5 percentage points from the previous best of 98.7% (Alibaba Qwen3.5 Omni Plus Realtime), with Flash variant scoring 96.1%. On Tau Voice, the Plus variant currently leads with a score of 54.6%, ahead of Grok Voice Think Fast 1.0 at 52.1%. On Full Duplex Bench, the Plus variant leads our Full Duplex Bench subset at 98.4%, with Flash at 96.9%, both ahead of the best non-Alibaba model, GPT-Realtime-2 (Minimal) at 96.1%. ➤ Speed: The Plus variant records an average Time to First Audio of 4.02 seconds on Big Bench Audio, with Flash at 4.16 seconds, among the slowest models on our leaderboard, and well behind GPT-Realtime-2 (Minimal) at 1.10 seconds and GPT-Realtime-2 (High) at 1.14 seconds ➤ Price: Plus costs $4.42 per hour of input audio on our Big Bench Audio subset, more expensive than GPT-Realtime-2 High ($4.14) and ~2.4x cheaper than GPT-Realtime-2.1 High ($10.75). The average cost for Flash variant is $4.77, higher than the Plus variant despite lower list prices, driven by comparatively more verbose responses. See below for more detail ⬇️
Show more
Open weights intelligence advanced today with the release of Kimi K3. The gap between the leading proprietary and open weights models is now just 4 points on the Artificial Analysis Intelligence Index, the smallest it has been since the GLM-5 release in February @Kimi_Moonshot's Kimi K3 scores 57 on the Artificial Analysis Intelligence Index. Only two labs have higher-scoring models: @AnthropicAI, with Claude Opus 5 at 61 and Claude Fable 5 at 60, and @OpenAI, with GPT-5.6 Sol at 59.
Show more
Compare Claude Opus 5 with other leading models at:
Full breakdown of the individual evaluations in the Artificial Analysis Intelligence Index for Claude Opus 5 with max effort
Claude Opus 5's effort setting spans a wide range of token usage-performance tradeoffs. On GDPval-AA v2, effort levels span 407 Elo points, with output token usage ranging around 8x from low to max effort. Like with GPT-5.6 Sol, this means Opus 5 can use either far fewer or far more tokens to complete the evaluation than models from other labs, depending on effort settings
Show more
Claude Opus 5 (max) is the new leader on both GDPval-AA v2 (1861 Elo, +114 over Claude Fable 5) and AA-Briefcase (1720 Elo, +146 over Fable 5). These benchmarks test the ability of models to produce accurate and well-presented professional outputs using our open source reference agent harness, Stirrup
Show more
Claude Opus 5 is narrowly the most intelligent model on the Artificial Analysis Intelligence Index, offering comparable intelligence to Fable 5 at 26% lower Cost per Task We supported @AnthropicAI to evaluate Claude Opus 5 ahead of release: it sets the highest GDPval-AA v2 and AA-Briefcase scores so far. Opus 5 (max) scores 61 on the Artificial Analysis Intelligence Index, effectively tied with Claude Fable 5 (max, 60), and ahead of GPT-5.6 Sol (max, 59), Kimi K3 (57), and Claude Opus 4.8 (max, 56) Key takeaways: ➤ New leader in agentic knowledge work: Claude Opus 5 (max) scores 1861 Elo on GDPval-AA v2, >100 points ahead of Claude Fable 5 and GPT-5.6 Sol (max). On AA-Briefcase, our proprietary agentic knowledge work benchmark, it scores 1720 Elo, +146 ahead of Fable 5. These benchmarks test the ability of models to produce accurate and well-presented professional outputs using our open source reference agent harness, Stirrup ➤ Joint first place on the Coding Agent Index: Claude Opus 5 (xhigh) with Claude Code leads the Artificial Analysis Coding Index, including the highest score on SWE-Atlas-QnA ➤ Frontier intelligence with reduced cost: Claude Opus 5 (max) costs $2.03 on average per Intelligence Index task, below Claude Fable 5 (with fallback) at $2.75, but still above Claude Opus 4.8 (max) at $1.80 and Claude Sonnet 5 (max) at $1.53. However, at high and xhigh reasoning efforts Opus 5 can outperform both Opus 4.8 and Claude Sonnet 5 at a lower cost per task ➤ Frontier agentic terminal use: 89% on Terminal-Bench v2.1 at max effort, roughly in line with the leader, GPT-5.6 Sol (xhigh) ➤ Outperformance on scientific reasoning: Along with leading agentic performance, Claude Opus 5 scores 53% on Humanity’s Last Exam in line with Fable 5; on CritPt, a frontier physics evaluation developed by Argonne and UIUC researchers, it also matches Fable 5 but sits behind GPT-5.6 Sol, GPT-5.5 Pro, and GPT-5.6 Terra ➤ Factual knowledge still lags Fable 5: As expected from the models’ size classes, Opus 5 still has lower factual knowledge on AA-Omniscience than Fable 5. It improves +7 points on AA-Omniscience Accuracy over Opus 4.8, but answers more often when uncertain - its hallucination rate rises +14 points to 50% ➤ Improving efficiency, but only on the Intelligence vs. Cost per Task Pareto frontier at high Intelligence levels: Opus 5 outperforms Fable 5 at lower cost, but at lower effort levels it sits just behind the GPT-5.6 family on the Intelligence vs. Cost per Task frontier Other model details: ➤ Context window: 1 million tokens (equivalent to Opus 4.8) ➤ Pricing: As with recent Opus launches, tokens cost $5/$25 per million tokens of input/output; cache pricing remains at a 25% premium for cache writes ($6.25 per million tokens) with 5-minute time to live, and 90% discount for cache hits ($0.50 per million tokens) ➤ Five effort settings (low, medium, high, xhigh, max), and support for server-side fallback as with Fable 5. Intelligence Index evaluations were run with Opus 4.8 fallback enabled
Show more
0
67
2.1K
210
Forward to community