Thanks Greg — ARC-AGI-3 is a remarkable benchmark. Few environments make it so clear that intelligence is not just about solving a static problem, but about learning through interaction over a long horizon.
This work actually started from a question I wrote down earlier this year: can an agent learn what to remember, update, or forget from environment feedback, and use that bounded internal state to make better future decisions? Here, “learning” happens through test-time memory updates rather than weight updates.
We’d be thrilled to have the ARC Prize evaluation team test dots3-note preview once the team is back, and would genuinely value the independent verification.