OpenAI has found new internal optimizations capable of cutting the cost of serving existing models by more than half. In other words, 50% cheaper inference or more, without swapping the model for a weaker one. On top of that, the engineers are developing the new chip, which also represents some incredible leaps in engineering.
All of this will start being implemented over the next few months. So we have a new family of models: Astra. We have a massive pretraining run with supposedly more than 10T parameters: Bel. And we have interesting, cheap technology behind all of it, making it possible for users to do twice as much, three times as much, or even more, for the same cost.
Meanwhile, I’m not seeing advancements nearly as significant from Anthropic, and even their “Model 2,” which is basically their ace in the hole, isn’t something they plan to release because it’s “too powerful and dangerous” (and also extremely expensive), so they prepared weaker checkpoints instead.