注册并分享邀请链接,可获得视频播放与邀请奖励。

GDP
@bookwormengr
AI model & hardware co-design, Inference economics Safe super intelligence for all All views strictly personal. No investment advice, dummies!
加入 July 2010
12.6K 正在关注    18.9K 粉丝
Your favourite benchmark @teortaxesTex !
Huge deal if true! CritPt is a difficult benchmark where top scores started around ~9–13% and began plateauing around 30–32% with only modest improvements each new model release. However, a new audit found benchmark errors in 21 of 56 questions. After repairing or removing them, GPT-5.6 Sol reached 94.4% in pass@4 on the corrected benchmark. The setup isn’t directly comparable to the official leaderboard, but it strongly suggests I overestimated the physics–math reasoning gap. Frontier models are probably (much) better at physics reasoning than I realized.
显示更多