Prime Intellect researcher
@sethkarten reveals the reason Prime Agent hit 95.5% on ARC-AGI-3 with Opus 5, matching the human expert baseline:
"The abstract and symbolic reasoning that is provided by learning how to model the world in these novel game environments in ARC-AGI-3 make it a very useful benchmark for general reasoning capabilities of the models."
"95.5% is a fantastic score. That was with Prime Agent running with Opus 5, and we found that Opus 5 was able to greatly take advantage of the capabilities of Prime Agent, more so than the other ones we benchmarked against."
"To put in perspective, the human baseline that ARC-AGI released regarding what they expect humans to be able to do in ARC-AGI-3 was 95.4. So we're on par now with potentially saturating the benchmark."
@a1zhang @PrimeIntellect