Register and share your invite link to earn from video plays and referrals.

Karim Mattar
@MattarARK
AI Research Associate @ARKInvest | AI, Cloud & Semis | Research x Automation | | Disclosure:
250 Following    6.8K Followers
A week ago OpenAI owned 11 of 15 seats on this frontier. Yesterday Opus 5.5 kicked Astra off the top and Anthropic owns more of the peak again, while Sol and Luna flood the cheap end at half the old price. This market resets every few days now. The labs that win are the ones who can live there.
Show more
How quickly things can change! Here's the updated frontier with Opus 5.5, GPT-6 Sol & Luna incorporated. Opus 5.5 has kicked Astra off the frontier, giving Anthropic more ownership of peak intelligence than before, while OpenAI still dominates the mid to low end of the frontier with GPT-6 Sol & Luna coming in smarter and 50% cheaper than their prior 5.6 generations. Xiaomi's MiMo takes the rotating open source spot for today, but otherwise the entire AA Pareto frontier is dominated by just Anthropic & OpenAI.
Show more
I've been turning effort dials like they're free upgrades. They're not. Anthropic gets most of the juice early. OpenAI keeps climbing if you pay for max. So the default seat and the burst seat might not be the same model anymore. That changes how you build agents more than another leaderboard screenshot.
Show more
There is something clearly different in how Anthropic & OpenAI scale effort levels. Doesn't show up on every benchmark, but these results on FrontierCode make it really clear. Anthropic models peak at lower effort levels, where as OpenAI models start low and climb up fairly consistently. The result is a better score at a lower cost for Anthropic models, but an unintuitive experience where increasing effort does not increase scores and might actually degrade performance (on this benchmark at least).
Show more
Sol and Luna basically kept 5.6-level intelligence and cut the bill in half. Opus 5.5 just did the same move from the other side and took the top of the index. If both labs keep selling smarter defaults cheaper every release cycle, the scarce thing stops being the model and starts being taste, eval harnesses, and who owns the workflow the agent never leaves.
Show more
GPT-6 Sol and Luna push the cost efficiency frontier by halving cost relative to GPT-5.6 Sol and Luna. Intelligence Index and Coding Agent Index scores remain level with GPT-5.6, with progress in some evaluations and regressions in others Pricing is approximately half that of GPT-5.6: Sol drops from $4/$20 to $2/$10 per million input/output tokens, and Luna from $0.20/$1.20 to $0.10/$0.50, with the same 90% discount for cache reads and 25% premium for cache writes. Key takeaways: ➤ Halves Cost per Task: GPT-6 Sol (max) costs $1.06 per task to run the Artificial Analysis Intelligence Index, ~50% less than GPT-5.6 Sol (max) at $1.99. GPT-6 Luna (max) costs $0.07 per task, ~60% less than GPT-5.6 Luna (max) at $0.18. This is driven by the price cut, as both models use slightly more output tokens per task (31k vs 29k for Sol, and 51k vs 41k for Luna). These two releases allow OpenAI to capture a significant portion of the cost efficiency Pareto frontier. ➤ In the Coding Agent Index, Sol improves but Luna regresses: In OpenAI's Codex harness, GPT-6 Sol (max) scores 57 in the Artificial Analysis Coding Agent Index, up 2 points from GPT-5.6 Sol (max), with gains in Terminal-Bench 4.0 (43% vs 37%) and SWE-Atlas-QnA (58% vs 54%). At $2.99 per task it costs ~50% less than GPT-5.6 Sol (max) and sits on the Pareto frontier of Coding Agent Index vs Cost per Task. GPT-6 Luna (max) scores 41, down 2 points from GPT-5.6 Luna (max), with lower scores in SWE-Atlas-QnA (44% vs 49%) and DeepSWE v1.1 (64% vs 66%), at ~60% lower cost per task. ➤ Significant reduction in hallucination: Both models hallucinate less in AA-Omniscience, our knowledge and hallucination benchmark. GPT-6 Sol (max) cuts its hallucination rate from 92% to 60% and GPT-6 Luna (max) from 93% to 77%. Sol achieves this by declining to answer more often: it attempts 83% of questions vs 99% for GPT-5.6 Sol (max), which cuts wrong answers by about a quarter but also lowers accuracy 5 points from 59% to 54%. Luna's accuracy is broadly unchanged at 44% vs 43% while it answers fewer questions. On the AA-Omniscience Index, Sol improves from 22 to 27 and Luna from -10 to 1. ➤ Mix of improvement and regression across evals: Beyond AA-Omniscience, both models improve in AutomationBench-AA (Sol 62% vs 60%, Luna 53% vs 50%) and Terminal-Bench 4.0 (Sol 44% vs 40%, Luna 13% vs 12%). However, we observe regressions in two key knowledge work evaluations. In GDPval-AA v2.1, our benchmark adapted from OpenAI's dataset of economically valuable tasks across 44 occupations, Sol drops ~100 Elo points and Luna ~75. Luna also drops ~45 Elo points in AA-Briefcase v1.1, while Sol is level. AA-Briefcase v1.1 is a private evaluation across multi-week knowledge work projects, with thousands of input files. Our team has manually inspected hundreds of model outputs: the regressions tend to be driven by reduced presentation quality and deliverables that omit rubric elements. Congratulations @OpenAI and @sama on the launch!
Show more
2 IN 1 DAYYYYY
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
Show more
GPT Sol 6 and my usage reset where are youuuuu ???
Use up all your credit, hit reset, continue like nothing happened lads
If you're on Pro, Max, or Team, your reset is available today in Settings → Usage. Apply it any time until Oct 22. Opus 5.5 is the default for paid plans. It's priced lower than Opus 5, so your 5-hour and weekly limits go 25% further.
Show more
Excited is an understatement
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
Grok 4.7’s coding agent score went from 47 to 56, putting Grok + Grok Build ahead of GPT-5.6 Sol in their native harnesses! Let's check out what those gains cost - Intelligence Index output token use more than doubled, but reasoning effort also increased from high to xhigh. If the extra reasoning means fewer failed tasks and less human cleanup, it could still be cheaper per completed job.
Show more
Grok 4.7 scores 46 on the Artificial Analysis Intelligence Index to bring SpaceXAI into the top 4 AI labs. Coding Agent Index performance has also improved, overtaking GPT-5.6 Sol Grok 4.7 scores +2 points over Grok 4.6 on the Intelligence Index, with strong performance on agentic knowledge work tasks. We evaluated the new model at xhigh reasoning effort. Congratulations to @SpaceXAI and @ElonMusk on the release! Key takeaways: ➤ Grok 4.7 joins the frontier of agentic knowledge work: Grok 4.7 gains +111 Elo over Grok 4.6 (high) on AA-Briefcase, our private benchmark for long-horizon agentic knowledge work, scoring 1657 Elo and placing it alongside Claude Opus 5 and Claude Fable 5.1 at the frontier. On GDPval-AA, it scores 1695 Elo, +90 ahead of Grok 4.6 (high). ➤ A leap in coding agent performance: Grok 4.7 (xhigh) with Grok Build scores 56 on the Artificial Analysis Coding Agent Index, up +9 points from Grok 4.6 (xhigh). Among models in their native harnesses, Grok 4.7 + Grok Build now ranks 4th, behind only Claude Fable 5.1, GPT-6 Astra, and Claude Opus 5. ➤ Incremental performance changes elsewhere: Outside of agentic knowledge work, Grok 4.7 broadly matches Grok 4.6 (high) on the other Intelligence Index tasks. It improves on Terminal-Bench 4.0 (+4.5 percentage points) and GDP.pdf (+3.0 p.p.), with regressions on AA-LCR (-3.7 p.p.) and AutomationBench-AA (-1.1 p.p.). ➤ High token use across tasks: Grok 4.7's gains come with higher token usage. Grok 4.7 (xhigh) uses approximately 81k output tokens per Intelligence Index task, compared with 36k for Grok 4.6 (high) and 27k for GPT-6 Astra (max) - 125% and 196% more, respectively. Other model details: ➤ Context window of 500k tokens, unchanged from Grok 4.6 ➤ Pricing of $2/$6 per 1M input/output tokens with cache hits discounted to $0.50 per 1M tokens, matching Grok 4.6 ➤ Configurable reasoning effort spans low to xhigh. Our evaluation uses xhigh.
Show more
When you look at open-weight models processed 56% of tokens on @vercel's AI Gateway in August, up from 13% in April. That’s 4.3x the share in four months. When you look even deeper you'll find they only accounted for 14% of estimated spend, while Anthropic still accounted for 64%. Seems like there’s room for a lot of cheap inference alongside models that can justify a premium. If more production workloads can run on the cheapest model that clears the quality threshold, application companies could keep more of the value they create. How much work becomes economical once inference costs stop being the constraint?
Show more
Trying out @higgsfield to create an ad video but can’t seem to get good output Anyone used it before or know some tips??
Muse is now available on Mac 💻. Your personal agent can get things done for you directly on your computer (all with your explicit permission). - Organize your downloads folder - Find a file you’ve lost track of - Summarize your messages and notes …more coming soon. Try Muse for Mac:
Show more
0
127
1.3K
109
Forward to community
Astra + Luna putting OpenAI on 11 of 15 Artificial Analysis Intelligence Index seats is the easy chart. The interesting constraint isn't "OpenAI is back." It's whether router/wallet share follows the pareto with a lag, or whether Anthropic keeps the production dollars while OpenAI owns the index. Falsifier: if OpenRouter/closed spend flips back to Anthropic while the index stays OpenAI-heavy, evals and wallets diverged. If wallet share tracks the 11/15, the index was the leading indicator.
Show more
With the combination of Astra & 5.6 Luna, OpenAI dominates the pareto frontier on the Artificial Analysis Intelligence Index, owning 11 out of 15 positions. This is why they are gaining share.
Show more
Anthropic and OpenAI are now hunting 20-30 MW deals on top of the GW campuses - UK, Nordics, US. The mechanism isn't "they ran out of ambition." It's that smaller allocations clear faster when the big builds slip, so inference and serving don't wait on a single campus COD. Falsifier: if the 20-30 MW hunt dries up once the next GW site unlocks, this was a stopgap. If it keeps expanding while Stargate/Nscale scale, the real market is a barbell of campus + edge MW.
Show more
Anthropic and OpenAI are now chasing 20 to 30 MW compute deals alongside gigawatt projects as “speed to usable capacity” becomes increasingly important. That puts already energized sites in a way stronger position as labs look to take whatever capacity can come online fastest.
Show more
Hi GPT 6 Astra one line of code will do and no you don’t need my SSN to fix RLS in Supabase thanks
This is actually a joke from Muse 🤣 Token cost is absurdly low even on the subscription Discretion - I am using contributor mode as I'm building something for myself, so will gladly take the 20x usage to share my prompts and codebase I'm on max effort, it's not as fast as Muse typically is but still impressive. Question for team @Meta - is there a difference in speed between contributor and standard mode?
Show more
$875M into Positron at ~$5B for Asimov - inference silicon built around LPDDR5X instead of HBM, with per-chip memory up into the TB class and a TSMC 3nm tapeout aimed at late 2026. The mechanism isn't "another AI chip." It's betting the serving bottleneck is memory bandwidth and HBM queue risk, so mobile DRAM + custom servers can undercut the HBM tax. Falsifier: if HBM supply loosens faster than Asimov ships, the LPDDR detour looks cute. If HBM stays tight into 27, memory-first is the real inference bid.
Show more
NEA recently led @positron_ai's $875M Series C — now @NeehaarGandhi @psikryan @lilatretikov @ScottDSandell & Forest Baskett are sharing why we think this is the next architectural reset for AI. ⚠️ AI just hit a trillion-dollar traffic jam — and it's not compute that's stuck, it's memory and networking. We think we found the company built to fix it. ⚡ Meet Asimov: a chip that ditches the switch entirely, links 16,000+ chips directly, and crams in more memory than anything else on the market. No HBM. No CoWoS. No bottleneck. 🚀 5x the tokens per dollar. A fraction of the power. And the its first-gen system already live.
Show more
Three open-weight labs grew monthly dollars spent on their models 10x or more on @OpenRouter so far in 2026. Closed models grew slower from much higher bases. The interesting constraint isn't "open is winning." It's that router spend is revealing preference under real prices - and the open stack is compounding from a small base while closed is still the wallet but not the growth curve. Falsifier: if closed models reclaim the growth rate into Q4 while open stays flat, this was a base-effect print. If open keeps 10x-class growth with rising absolute dollars, the mix shift is real.
Show more
Three key open-weight labs have seen monthly dollars spent on their models grow by 10x or more so far in 2026. The closed models have seen comparatively slower growth in spending, though from much higher starting points.
Show more
Yep there she goes Genuinely used Fable for 2 sessions man 😭 Don't get me started on Astra consumption on the subscriptions...you'd be lucky to get half a sesh in. Efficiency only really applies when you're using the API
Show more
I am definitely feeling this usage reduction on the @claudeai subscriptions... Fable for 2 sessions (PM in one of them, not even builder) and my whole weekly limit is gone
Muse has officially earned its place as one of my coding agents. Well deserved mate, the token value is absurd And if you let them use your messages and codebase to improve the model, you get a 20x discount 😭
Show more