Huge deal if true! CritPt is a difficult benchmark where top scores started around ~9–13% and began plateauing around 30–32% with only modest improvements each new model release. However, a new audit found benchmark errors in 21 of 56 questions. After repairing or removing them, GPT-5.6 Sol reached 94.4% in pass
@4 on the corrected benchmark. The setup isn’t directly comparable to the official leaderboard, but it strongly suggests I overestimated the physics–math reasoning gap. Frontier models are probably (much) better at physics reasoning than I realized.