Bixbench3: Can models reproduce the entire analysis from papers? We had agents repro papers end-to-end, with over 1B tokens spent on some. Frontier agents still score below 50%, and due to real problems like environment misconfig, quitting, or faking data 1/6