注册并分享邀请链接,可获得视频播放与邀请奖励。

sarah
@sarahtonin
solve physics - my views evolve - optimist
加入 July 2012
2K 正在关注    12K 粉丝
Huge deal if true! CritPt is a difficult benchmark where top scores started around ~9–13% and began plateauing around 30–32% with only modest improvements each new model release. However, a new audit found benchmark errors in 21 of 56 questions. After repairing or removing them, GPT-5.6 Sol reached 94.4% in pass@4 on the corrected benchmark. The setup isn’t directly comparable to the official leaderboard, but it strongly suggests I overestimated the physics–math reasoning gap. Frontier models are probably (much) better at physics reasoning than I realized.
显示更多
0
5
108
9
转发到社区