Thanks Greg — ARC-AGI-3 is a remarkable benchmark. Few environments make it so clear that intelligence is not just about solving a static problem, but about learning through interaction over a long horizon.
This work actually started from a question I wrote down earlier this year: can an agent learn what to remember, update, or forget from environment feedback, and use that bounded internal state to make better future decisions? Here, “learning” happens through test-time memory updates rather than weight updates.
We’d be thrilled to have the ARC Prize evaluation team test dots3-note preview once the team is back, and would genuinely value the independent verification.
1/6 42 has been in my handle for years - picked it on a whim. Never thought the dots model would one day achieve exactly that: 42/42 at IMO 2026.
Officially certified full marks. Gold-medal level. Perfecto! Still feels surreal.
A few thoughts on why this matters - and curious what others think.