Just sitting and reflecting on this GPT-6 Astra graph.
GPT-6 Astra set a new SOTA on ARC-AGI.
๐งฉ ARC-AGI-1 โ 98.5%
๐ง ARC-AGI-2 โ 95.0%
๐ค ARC-AGI-3 โ 99.95%
But the ARC-AGI-3 result has a weird twist.
ARC-AGI-3 drops Astra into a completely NEW interactive environment with no instructions.
It has to ...
๐ explore
๐ง figure out the rules
๐บ๏ธ build a world model
๐ remember what it learned
๐ฏ develop a strategy
๐ adapt when it's wrong
ARC tested ~500 humans to establish how efficiently people solve these environments.
GPT-6 Astra with ARC's standard agent harness:
62.7%
Astra using the new Provider Adapter that preserves reasoning state between requests + compacts it:
99.95% ๐คฏ
Give a powerful model a persistent working state + good compaction, and suddenly the same intelligence can behave VERY differently over a long task.
That's going to matter for Local AI too.