"One remaining area of important capability we haven't seen demonstrated yet by any model is open-ended invention and discovery."
Benchmarks keep going up. Open-endedness keeps moving up in importance.
GPT-6 Astra is the new SOTA on ARC-AGI-3
It's a qualitatively large leap towards AGI and the pace of progress is frankly surprising. That said, we lack evidence to call this AGI yet.
While we are still studying the human capability gaps, we believe open-ended invention is unsolved, and this will form the new basis for ARC-AGI-4.
Reminder that ARC v3 tests whether models can figure out how to efficiently make sense of unfamiliar environments and operate autonomously toward goals within them.
Summary: properly equipped Astra can now do this.
Astra direct model scored 66% at ~$500/game. Up from Sol 8%. This is apples-to-apples with all other ARC v3 scores we've verified to date. This score gives the best look at the general intelligence of base Astra.
Using an OpenAI-specific adapter (which is open source) we verified the first ~100% at $300/game. If you're using Astra via API I would prioritize migrating to /conversations endpoint with compaction enabled.
Direct model testing continues to carry important scientific interest towards AGI. But provider adapters are a better estimate of what you should expect day-to-day when using these models inside of products.
The action efficiency of Astra is also pretty incredible. With the provider adapter, Astra used 50% fewer actions on average than our human baseline (who both score 100%).
A surprise for us is that across all models, action efficiency on v3 is very bi-modal. Models either get the game or not. Weaker agent models aren't able to "brute force" their way to an inefficient understanding.
Tactically, our results support always using Astra's higher reasoning tiers when deploying agents into environments. It not only makes Astra better but is cheaper too (as a result of needing less turns).
The key technique Astra seems to leverage is on-the-fly symbolic world modeling. It develops a custom algebra and notation for the environment through interaction and uses this understanding to plan (all in token space).
This approach is similar to earlier harnesses we saw which created symbolic models through side-car executable programs to model the environment and plan.
Based on Astra results I now see clear paths to recursive self improvement via horizontal data scaling and continual learning where the learned state is information stored and managed outside of the model weights (text, programs, databases, etc).
I expect more and more "learning" to happen outside the model weights and this will be of growing interest to both researcher and AI product builders.
One remaining area of important capability we haven't seen demonstrated yet by any model is open-ended invention and discovery.
Humans clearly posses this capability and we've so far seen zero evidence that frontier AI -- even as powerful as Astra -- is capable of this feat.
Consider AI research 10 years ago vs today. Humans invented so much: transformer, self-supervised LLM pretraining, diffusion, RLHF, test time adaptation, ...
Society needs AI systems that not only can autonomously operate towards defined goals. We need AI systems help craft great futures that are inherently undefined.
I believe we in at a very tenuous spot. We now have extremely powerful automation system that capable of causing a lot of disruption.
Without the counterbalancing force of demonstrated invention by AI, I fear society will turn more and more against open frontier AI research and development -- trapping us in a very bad middle ground.
We must move as swiftly and with focus towards this important AI capability. And this has become the primary focus of our work on ARC-AGI.