We ran the largest open experiment on how frontier models do AI research.
100+ autonomous runs across 10+ models, sandboxed on 8xH200s for up to 8 days, iterating on the nanoGPT optimizer track.
Best runs closed 82% of the gap to a record built by dozens of humans over months.