Just sitting and reflecting on this GPT-6 Astra graph.
GPT-6 Astra set a new SOTA on ARC-AGI.
🧩 ARC-AGI-1 → 98.5%
🧠 ARC-AGI-2 → 95.0%
🤖 ARC-AGI-3 → 99.95%
But the ARC-AGI-3 result has a weird twist.
ARC-AGI-3 drops Astra into a completely NEW interactive environment with no instructions.
It has to ...
👀 explore
🧠 figure out the rules
🗺️ build a world model
📌 remember what it learned
🎯 develop a strategy
🔄 adapt when it's wrong
ARC tested ~500 humans to establish how efficiently people solve these environments.
GPT-6 Astra with ARC's standard agent harness:
62.7%
Astra using the new Provider Adapter that preserves reasoning state between requests + compacts it:
99.95% 🤯
Give a powerful model a persistent working state + good compaction, and suddenly the same intelligence can behave VERY differently over a long task.
That's going to matter for Local AI too.