Register and share your invite link to earn from video plays and referrals.

sarah
@sarahtonin
solve physics - my views evolve - optimist
2K Following    12K Followers
Huge deal if true! CritPt is a difficult benchmark where top scores started around ~9–13% and began plateauing around 30–32% with only modest improvements each new model release. However, a new audit found benchmark errors in 21 of 56 questions. After repairing or removing them, GPT-5.6 Sol reached 94.4% in pass@4 on the corrected benchmark. The setup isn’t directly comparable to the official leaderboard, but it strongly suggests I overestimated the physics–math reasoning gap. Frontier models are probably (much) better at physics reasoning than I realized.
Show more
not to be a buzz kill but “find a counter example” always seemed like the lowest hanging fruit for AI assistance in math research.. this is cool but if it’s a huge update for you then you probably were just really poorly calibrated
Show more