ICLR 2027 submission: we benchmarked all frontier models and our model performs...
Anthropic & OpenAI today: no you don't.
ICLR deadline is 9/25, so i guess no excuse not including opus5.5 and gpt6-luna & sol in the paper.
Opus 5.5 apparently built with a weather system that matches whatever weather is in NYC at the time...
So with the nor'easter, there's rain, and the NPCs are carrying umbrellas!
It's the little things :)
Opus 5.5 is 20% cheaper per input and output token than Opus 5, and 60% cheaper on cache reads. So what does that actually do to the cost of a task in Claude Code?
We ran the numbers, and built a calculator so you can run yours from /usage:
opus 5.5 really changed my expectations for the next openai release.
before this i would’ve assumed 6.1 astra would comfortably move things forward again.
now i’m not nearly as convinced.
especially because anthropic still has fable 5.5 coming.
openai better have something fucking ridiculous waiting.
Opus 5.5 just filed a pull request that changes the streaming behavior in T3 code.
It's so cool that I can send off a prompt with a vague idea of what I want, and the result is a ready-to-merge pull request with a video demo of the changes directly in the PR.
The PRs I make with AI are significantly better than the ones I used to make by hand. The future is awesome.
opus 5.5 still isn't better than gpt-6 astra at xhigh/max in raw intelligence
for STEM especially, i'd still take gpt-6 pro pretty comfortably. opus 5.5 probably has the edge in coding, but it's a narrower advantage
astra is also still unchallenged in computer use and vision