Claude Sonnet 5.5 is out! We wrote a guide for building with it:
• choosing between Sonnet 5.5 and Opus 5.5
• migrating from Sonnet 5 and tuning effort
• using it in Claude Code
Claude can now help you build evaluations and hillclimb on them.
In this article, we share guidance on eval design & skills that Claude Code can use to improve your applications.
Claude Sonnet 5.5 (max) makes large strides on Terminal-Bench, sitting among the top models for both Terminal-Bench 4.0 and Terminal-Bench-Science. In Terminal-Bench 4.0 it scores 64%, a 50 point increase over Claude Sonnet 5 (max), and slightly above 60% for Opus 5.5 and GPT-6 Astra (xhigh).
On our leaderboard for Terminal-Bench-Science - a benchmark of agentic terminal use to complete realistic scientific research workflows across domains - it scores 53% and sits behind only GPT-6 Astra and Opus 5.5. Terminal-Bench-Science is not currently included in the Artificial Analysis Intelligence Index.
Claude went from the subscription that lasted the least to the one that lasts the most. OpenAI went the other way, right around DevDay and the launch of their $500 subscription.
Kinda weird timing to make the limits feel tighter...
Claude's upgrade pages, where free users see the Pro, Max, and Team plans and paying users move tiers, drew 21.0M visits from 14.7M users in Aug-26 (+400% and +476% Y/Y).
Read more about Claude's performance ahead of Anthropic's expected IPO in our full report:
Claude's weekly limit reset is in!
Seeing Codex fanboys celebrating Tibo initiating a limit reset for them, so why not do mine by with Anthropic's chatbot :)
I'm gonna be trying different stuff with Opus 5.5 and will probably post it all here while having fun.
Crazy how much value you can get from just a $20/$100/$200 monthly plan.
Claude Opus 5.5 feels like the first model to have truly solved video animation.
We one-shotted this promo video for @dubdotco using Opus 5.5 High via @cursor_ai and it worked remarkably well 🤯
h/t @pedrooladeira for the 🔥 vid
Claude just computed a nine-loop scattering amplitude in planar N=4 super-Yang-Mills, past the eight-loop record.
What that means. Scattering amplitudes predict how particles behave & physicists compute them in layers of correction called loops. Each loop makes the answer more precise & costs exponentially, sometimes factorially, it’s more work.
Most real amplitudes have only been taken to two or three loops. The most precise prediction in particle physics, the electron's anomalous magnetic moment, used five.
The challenge came from physicist Matt von Hippel, who publicly asked whether an AI could push past eight loops using only compute an academic could afford.
Given one prompt & periodic instructions to continue, Claude ran largely unsupervised for days & did it, using bootstrap methods that Lance Dixon & collaborators developed.
Total cost: a few thousand dollars.
Dixon, who held the prior record, checked the result independently.
Three caveats that matter more than the headline: 1. N=4 super-Yang-Mills is a toy model which is a testing ground for techniques, not a theory describing our universe. No new real-world particle prediction came out of this. 2. Claude applied existing human methods rather than inventing new ones. 3. A human-led group at the Chinese Academy of Sciences, using GPT-6 for parts of it, had concurrently obtained most of the same result & published a dataset on September 17.
So the honest version isn't that a machine beats humans. It's that a hard calculation at the edge of a field now costs only a few thousand dollars & takes several days, by more than one route at once.