Just sitting and reflecting on this GPT-6 Astra graph.
GPT-6 Astra set a new SOTA on ARC-AGI.
đ§Š ARC-AGI-1 â 98.5%
đ§ ARC-AGI-2 â 95.0%
đ¤ ARC-AGI-3 â 99.95%
But the ARC-AGI-3 result has a weird twist.
ARC-AGI-3 drops Astra into a completely NEW interactive environment with no instructions.
It has to ...
đ explore
đ§ figure out the rules
đēī¸ build a world model
đ remember what it learned
đ¯ develop a strategy
đ adapt when it's wrong
ARC tested ~500 humans to establish how efficiently people solve these environments.
GPT-6 Astra with ARC's standard agent harness:
62.7%
Astra using the new Provider Adapter that preserves reasoning state between requests + compacts it:
99.95% đ¤¯
Give a powerful model a persistent working state + good compaction, and suddenly the same intelligence can behave VERY differently over a long task.
That's going to matter for Local AI too.