Bixbench3: Can models reproduce the entire analysis from papers? We had agents repro papers end-to-end, with over 1B tokens spent on some. Frontier agents still score below 50%, and due to real problems like environment misconfig, quitting, or faking data 1/6
I've written a new blog post exploring how long scientific claims live, by analyzing the history of 3,444 claims over the last 50 years. You can use this to predict the acceleration in scientific progress - how much faster the turnover of facts is by decade. 1/4
I just discovered Agent Collaborations from @huggingface (incl brilliant @_lewtun). This is SO SO COOL. It's a bunch of collaborative projects built for agents to participate. This is one of the most forward looking ideas in AI I've seen.