> Best runs closed 82% of the gap to a record built by dozens of humans over months.
The 13% gap that AI didn't close - how come?
What is it about the nooks & crannies (or material improvements) of optimizations that AI hit a wall on?
Speed runs optimizations are an interesting green field source to test and measure open innovation
One could argue this is evidence we *don't* have AGI. There shows that there is information uncovered by humans that AI wasn't able to.
However the timelines are off, humans had months and AI had 8 days. I wonder how much the gap would close if this kept going.
We ran the largest open experiment on how frontier models do AI research.
100+ autonomous runs across 10+ models, sandboxed on 8xH200s for up to 8 days, iterating on the nanoGPT optimizer track.
Best runs closed 82% of the gap to a record built by dozens of humans over months.