I am really enjoying these *CRUXes*. I was definitely more bullish on agents' prospects for writing a neurips-worthy paper than most, but that was largely because of, ahem, a certain scepticism about that as a bar. Still, this was definitely an update for me, as of now—but like Sayash's earlier benchmark, CoRE, I think this one will fall pretty soon.
My more general as to how is that we need to expand this out to way more kinds of papers—after all, some papers require a tonne of insight and creativity; but some pretty useful papers are really about just moving the needle forward a little bit. I'd like to see if there's somewhere in that distribution that the models can reach, and where they tap out. It would also make for a pretty sick burn for papers that are so unimaginative that they could have been written by one of today's frontier AI models...