Continual Harness: An Efficient Self-Improving Agent on ARC-AGI-3 by
@sethkarten from
@PrimeIntellect
> The heavy test-time learning required by the benchmark (ARC-AGI-3) pushes agents to form an internal world model of the rules and mechanics that updates with new evidence.