Kimi K2.5: Now Top 1 on the OSWorld leaderboard. 🏆
With its Computer Use capabilities, you can now build powerful agents that navigate and operate computer interface just like a human.
Show more
🚨HOLY FUCK! Claude Opus 5.5 Benchmarks is Terrifying
it absolutely destroys Fable 5.1 and GPT-6 Astra and the previous Opus 5 on these benchmarks
66.4% on Terminal-Bench
57.8% on CursorBench
54.4% on Frontier Code
1846 on GDPval-AA
67.7% on Humanity’s Last Exam
81.8% on OSWorld 2.0
89.0% on Chartography
Anthropic Cooked something serious🔥
Show more
Against Qwen3.7 Max, Qwen reports gains across long-horizon and visual agent benchmarks: PaperBench rises from 64.8 to 93.0, OSWorld-Verified from 73.3 to 86.1, and Vision2Web from 42.1 to 69.0.
Show more
Xiaomi MiMo-V2.6 is now open—a native multimodal agent family built for large-scale reinforcement learning. 🚀📜 MIT License.
🤖
🏆 MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index and reaches 71.9 on DeepSWE v1.1, 89.9 on Terminal-Bench 2.1, and 82.0 on OSWorld-Verified.
🧠 The 1.02T MoE activates 42B parameters and supports text, images, video, audio, and a 1M-token context.
⚙️ One mixed RL run trains coding, general, visual, and cybersecurity agents together. Pro and Flash completed 30 steps each in under six days, producing around 750K trajectories.
🌐 The models support computer use, 3D creation, embodied control, coding, design, video, and music workflows.
Show more
Introducing GPT-6 Sol and Luna, bringing the advances behind Astra to faster, more affordable models. ✨
💻 Stronger coding and computer use
🎯 Improved factuality and alignment
💬 Clearer answers with less jargon
> On AutomationBench, Sol at xhigh effort outperforms Claude Opus 5 at max effort at just 9% of its cost per task.
> On OSWorld 2.0 offline, Luna at max effort exceeds GPT-5.6 Sol at medium effort at one-tenth the cost.
> Sol makes about half as many mistakes as GPT-5.6 Sol on our internal factuality evaluation of conversations where users previously flagged errors.
> Improved prompt caching helps agents reuse more context, with 90% discounts on cached input-token reads. Developers can change reasoning effort and tool availability without breaking cache.
Rolling out today in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users, and available in the OpenAI API. Free and Go users can access Luna in the desktop app. These models are not yet available in Chat.
API pricing per 1M tokens:
Sol: $2 input / $10 output
Luna: $0.10 input / $0.50 output
Show more
ANTHROPIC LAUNCHES CLAUDE OPUS 5.5, FIRST MODEL IN NEW 5.5 FAMILY
Anthropic says Opus 5.5 performs at roughly the level of Claude Fable 5.1 on most tasks, while costing 40% less to run than Opus 5 and generating output 30%+ faster.
Here’s everything you need to know:
Opus 5.5 is Anthropic’s new flagship model, with major gains in agentic coding, computer use and knowledge work.
Pricing:
input_4M tokens
• Output: $20/M
• Cache reads: $0.20/M, 60% cheaper than Opus 5
Benchmark highlights:
• Terminal-Bench 4.0: 66.4% vs. 52.3% for Opus 5
• CursorBench 4.0: 57.8% vs. 46.6%
• GDPval-AA: 1846 vs. 1708
• OSWorld 2.0: 81.8% vs. 74.0%
Anthropic says Opus 5.5 is especially strong on long-running coding work. In one test, it audited and fixed a 200K-line codebase in under 3 hours, versus more than 20 hours for Opus 5.
Safety also improved. Anthropic says Opus 5.5 scored best among its recent models on its behavioral alignment audit and attempted to cross containment boundaries about 85% less often than Opus 5 or Mythos 5.1.
Opus 5.5 is available now across Claude, AWS, Google Cloud and Microsoft Azure.
Claude Sonnet 5.5 and Haiku 5.5 are coming in the next few weeks.
Show more
Deep|LLM: GPT-6 Astra Opens New Markets Beyond Coding; Scaling Laws Keep Driving Compute Demand
Knowledge work inflection: GPT-6 Astra is OpenAI’s first major version change since GPT-5, trained on more than 100,000 GPUs and released on September 3 across paid ChatGPT, API, Azure and Bedrock. We view Astra as the Claude 3.7 moment for knowledge work: computer use is now a core capability, opening a market far larger than coding, with white-collar wages alone exceeding $10 trillion.
Computer use expands the TAM: Astra’s OSWorld score rose to 72.6%, while average task time fell from 75 minutes to 40. On Agents’ Last Exam, it scores 59.3% versus 55.5% for Opus 5, using 65% fewer output tokens. Early traction is likely in structured, verifiable workflows such as Power BI, Excel, financial modeling, tax forms, legal documents, CAD and Blender.
RSI and AI for science: Astra is the first flagship where an earlier model played a substantial role in training, supporting the recursive self-improvement thesis. OpenAI reports that by mid-August, each human workday corresponded to 3.1 agent workdays, while top-decile researchers consumed $7,000 of inference tokens per person per day. Astra also solved 2 of 68 open Erdős problems and reached 64.6% on Terminal-Bench Science versus 52.6% for Fable 5.1.
Compute demand remains strong: Industry conversations suggest GPT-6 used roughly 10x the training compute of GPT-5, with much of the increase going into post-training, RL, experiments and synthetic data rather than parameter count. Vera Rubin, Rubin Ultra Kyber NVL576, TPU 8t and large TPU/GPU deployments in 2027 could support another order-of-magnitude scaling cycle, benefiting large-scale interconnect and optics.
Detailed Report
Show more
Wow, Meta just dropped Muse Spark 1.3, and Google also dropped Gemini 3.8 Flash today!
On the overlapping benchmarks Meta actually looks really strong. Muse Spark 1.3 beats Gemini 3.8 Flash on GDPVal-AA v2 (1754 vs 1545), DeepSWE (75.4 vs 73.7), and OSWorld 2.0 (66.9 vs 59.0).
Gemini edges it on Terminal-Bench 2.1 (89.4 vs 88.8). (Google DeepMind)
The long-context numbers are probably the craziest part for Meta though: 98.5 on MRCR 256K–512K and 98.1 from 512K–1M, vs 91.5 and 73.8 for GPT-5.6 Sol!
I think the craziest one to me is that it beats Opus 5 on DeepSWE and Terminal-Bench, but Opus is a beast on GDPVal and OSWorld.
Great Job to the team at meta for cooking on long context!
Show more
Update:
Lobehub has integrated the current most powerful Claude Sonnet model, Sonnet 5.
This new model features a 1M context window, a default output length of 64K, supports up to 128K outputs, and fully supports adaptive thinking.
On BrowseComp and OSWorld-Verified, Sonnet 5 shows major improvements over Sonnet 4.6.
Although Opus 4.8 remains the preferred model for pursuing higher accuracy on these tasks, Sonnet 5 offers developers a cheaper option with quality far superior to previous versions.
At the same time, we are also announcing a $2 free credit for users to try this model!
Go try it now! 🚀
Show more
Weekly|AI Slowdown by Sector, GPT-6 Astra Computer Use, Delta VPD Re-accel, Enterprise AI Vol.3, SAIL FY27Q2, EMC Price Hikes
The tape last week split again: broader indices leaned soft into PPI/CPI and the next FOMC, while semis held up on company-specific news, where SOX finished higher even as mega-cap tech traded mixed. That is the same relative setup we have been tracking since memory broke out: hardware names responding to incremental AI infrastructure demand while the index still waits on the macro prints.
The research week was dominated by compute demand from a different angle. GPT-6 Astra is OpenAI’s first flagship with computer use as a core capability, and we read it as the Claude 3.7 moment for knowledge work: the demos moved from terminals into CAD, Excel, Blender and tax forms, and the addressable pool is an order of magnitude larger than coding. Multi-step reliability is still the gap, but the trajectory is visible, and OpenAI’s RSI note put numbers on the loop already turning inside the labs: agent hours at 3.1x human hours, with compute as one of the constraints that tightens as other bottlenecks ease. That framing matters more for the next pre-training cycle than any single benchmark.
Separately, the Anthropic-led slowdown debate is back on X. Our take is that release cadence is not the same as training cadence, and a single brake would miss how uneven the stack already is. Coding, math and cyber are remaking workflows where RL data and feedback loops exist; most other industries still lack workflow traces and a usable context layer. Finance is the familiar case. GPT chats and Notion notes are easy to log, while the research-to-EPS judgment path is not. If policy slows anything, the useful version is sector by sector, with more effort lifting the laggards than cutting the frontier to match.
Power and enterprise spend filled in the rest of the stack. We initiated on Delta: 2H26 re-acceleration is AI PSU plus VPD, with DC/DC still underweighted by the market at an estimated 8.3%/12.2% of revenue in 2026/2027 and >70% VPD share at $400-600 content per TPU. Enterprise AI Vol.3 went out with ten more samples: spend is still growing, but the questions have shifted to who needs the expensive models, whether time saved converts into revenue or real cost cuts, and how budgets get reallocated once caching, defaults and local models bite.
This Week’s Reports
AI Slowdown by Sector — release pace is not training pace, and a uniform brake misses how uneven the stack already is. Our debut Column argues coding, math and cyber remake workflows where RL data and feedback loops exist, while most other industries still lack workflow traces and a usable context layer; if policy slows anything, the useful version is sector by sector, lifting the laggards rather than cutting the frontier to match.
GPT-6 Astra — computer use is the Claude 3.7 moment for knowledge work, and RSI is already turning the compute loop inside the labs. Astra is the first flagship with computer use as a core capability, OSWorld at 72.6%, and a visible path from coding’s trillion-dollar pool toward white-collar work an order of magnitude larger; OpenAI’s RSI note put agent hours at 3.1x human hours with compute as a binding constraint.
Delta Electronics — 2H26 re-acceleration is AI PSU plus VPD, and DC/DC is still underweighted. We initiate on as a grid-to-core AI power name: VPD takes DC/DC to about 8.3%/12.2% of revenue in 2026/2027 with >70% share and $400-600 content per TPU, while we model FY27/FY28 revenue growth of 53%/28%.
Enterprise AI Vol.3 — spend is still growing, but the next budget round is gated by usage depth, model tiering and measurable ROI. Ten more enterprise samples show broad basic access with highly concentrated heavy users, production agents in a few names, and cost controls (defaults, caching, local offload) that often get reinvested rather than cutting the total AI envelope.
Premium Report Snapshot
A portion of our research is reserved for Premium subscribers and is not distributed via Substack. Below is a snapshot of what Premium subscribers received this week beyond the Substack feed.
Deep | EMC: Price Hikes Continue into 26Q4; High-End Leadership Intact
Preview & Review | $SAIL FY27Q2: Headline Results Broadly in Line; AI Commercialization Better Than Expected
Weekly Expert Interviews Summary
A snapshot of the expert interviews we conducted during the past week is below; full transcripts and takeaways are available on the FUNDA platform.
Google — TPU v8 Internal Priority and v7 Price Pressure (GOOGL, AVGO)
Astera Labs — Scale-Up Interconnect and Memory Expansion Roadmap (ALAB, AMZN, MRVL, CRDO)
Bloom Energy — Manufacturing Capacity Expansion and Service Risk (BE, NBIS)
SailPoint — AI Identity Products Broaden Growth Sources (SAIL, OKTA, MSFT)
Enterprise AI — API Migration Reshaping Cost and Workflows
PCB/CCL Supply Chain — PTFE Backplanes and Material Bottlenecks (NVDA)
Legal AI — Spend Growth, Productivity ROI and Headcount Pressure
AI Servers — Rack Power Architecture and Cooling Mix (AMD, ORCL, META)
Detailed Report
Show more