가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

GDP
@bookwormengr
AI model & hardware co-design, Inference economics Safe super intelligence for all All views strictly personal. No investment advice, dummies!
가입 July 2010
12.6K 팔로잉 중    18.9K 팬
Your favourite benchmark @teortaxesTex !
Huge deal if true! CritPt is a difficult benchmark where top scores started around ~9–13% and began plateauing around 30–32% with only modest improvements each new model release. However, a new audit found benchmark errors in 21 of 56 questions. After repairing or removing them, GPT-5.6 Sol reached 94.4% in pass@4 on the corrected benchmark. The setup isn’t directly comparable to the official leaderboard, but it strongly suggests I overestimated the physics–math reasoning gap. Frontier models are probably (much) better at physics reasoning than I realized.
더 보기