Opus 5.5 is noticeably better than Opus 5.
But GPT-6 Sol isn't noticeably better than GPT-5.6 Sol.
So I ran them head to head on 10 real use cases.
Worth a 3 min read.
GPT-6 Sol and Luna just landed in Astra’s orbit.
Both launch today with API prices 50% lower than GPT-5.6.
Build with Sol. Scale with Luna. To production and beyond.
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
Grok 4.5 was an incredible model for the price: fast, pleasant to use, reliable, solid default model.
Grok 4.6 was a (forgivable) step in the wrong direction IMO: slower and more expensive, using way more tokens per task for a slight edge in intelligence. I get it, though. They have to climb benchmarks.
Grok 4.7 is much harder to forgive. They claimed it would be more token-efficient, and it's less by 30 to 80%. It scores worse than Grok 4.6 in various benchmarks. It's slower, it's less pleasant to use, and real-world costs come out to more than 2x above Grok 4.6, putting it over Astra's costs in real-world use.
Considering how much they've been hyping this model release up, I have to say I'm disappointed. The benchmarks don't tell the whole story, and it is pleasant to use in various real-world engineering tasks, but it feels so 2025 still.
The frontend capabilities are unacceptably bad. The 3D capabilities are nonexistent. It gets stuck in random Gemini-style loops all the time.
This was a very disappointing release. I hope that the SpaceXAI team can acknowledge that and impress us with the next one.
Whether it’s 3% or 4% I think the growth comes from people and businesses producing more with the same headcount.
Early movers will capture most of it before it becomes the baseline, so figure out right now where AI 10x’s your own output.
a week after launch, muse is now the #1# app in the App Store!
it’s been SO exciting to see how much y’all have done with muse. we can’t wait to do more together!
Honestly, this narrative makes sense. There may be an incentive.
But “incentive exists” doesn’t mean “crisis is fake”.
I haven’t seen Eisman say anything publicly about recent AI incidents and/or safety/alignment research.
Big Short investor Steve Eisman says the AI labs know they have no moats and are manufacturing a crisis to get regulation that hands them a duopoly
"I think this is all nonsense."
[ You think it's all nonsense? ]
"All nonsense. I think that there's something else completely going on here. ... What I think is happening is that token maxing is over. The open weight models are taking big market share. I think these companies are very nervous. They realize that there are no moats around their business whatsoever, and they're trying to manufacture a crisis that will create regulation, and that they think they can then manipulate to create the moats, to create the duopoly that they want."
[ Wow. ]
"That's what I think is going on."
"Honestly, I think this whole Terminator thing is garbage. That's for sure."
The gap is so wide it blows my mind.
The average person opens up ChatGPT a few times a month for a "Google search" or to make a funny image.
Meanwhile, we're having national debates about banning superintelligence and human extinction.
Almost nobody pays for AI out of pocket yet
Only ~3% of US consumers do, up from under 1% in 2023
The youngest buyers are adopting at 4x the rate of the oldest
Charts of the Week:
Dario Amodei talked about what a one-person billion-dollar company could look like.
He didn't give an official checklist, but I turned the examples from his answer into three filters to choose a business I could realistically build.
→ Does the business deploy its own capital?
Dario's example was a proprietary trading firm.
A real estate flipping business or used car dealership follows a similar pattern: use your money to buy something, improve or trade it, and sell it for more.
You may not need thousands of customers, but you do need capital, expertise, and the ability to absorb losses.
I ruled this route out because I wanted something someone could start without risking a large amount of their own money.
→ Can software deliver the result?
Claude Code makes building useful software much more accessible.
But I could still build myself a full-time job if every customer needs custom onboarding, a sales call, and ongoing help.
I wanted a result the software could deliver repeatedly.
→ Can sales and routine support be automated without making the customer experience worse?
This narrowed the options the most.
The customer needs to understand what they're buying, get started easily, and receive value without a consultation.
A file converter fits that pattern.
A product that requires me to explain a different implementation to every customer probably doesn't.
I landed on Agent Report Card: software that tests customer support agents and produces a report showing where they fail.
An agency can connect its agent, run the tests, and get a deliverable it can share with its client.
Routine questions can go through AI support, while refunds and account deletion still come to me.
That's the kind of workload I was trying to design for one person.
Astra for Law: Frontier intelligence built for your practice.
A new offering powered by GPT-6 Astra with tools, settings, and context to support the expertise and judgment of lawyers and legal technology firms.
AI systems are getting more powerful, and they're increasingly being used to build the next version of themselves. We want to illuminate that progress for the public.
Today, we're sharing three measurements that help track AI development:
1. How much AI R&D is done by AI.
2. How well AI agents are overseen.
3. How compute is allocated.
We provide a snapshot of these metrics from inside Anthropic. Any frontier developer could publish the same measures, and third parties could verify them.
As the world considers pacing the frontier, we should do everything possible to minimize the gap between what frontier labs know and what the public knows. This means better measuring the development of AI, publishing our findings, and giving society an opportunity to decide how to use this information.
Read the full post and methodology:
Higgsfield's NEW API finally allows you to pay per generation rather than a monthly rate with credits.
It isn't automatically cheaper though...
It's all about finding the breakeven point.