註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Zeyi(Andy) Liu
@ZeyiAndyLiu
PhD student at Yale University MMath IQC Waterloo/Perimeter Institute | BA Cambridge Now: Machine Learning Prev: Quantum Error Correction
加入 September 2023
479 正在關注    298 粉絲
New paper: How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks Frontier models have tackled Navier–Stokes. People naturally wonder, how is their capability for physical sciences? According to today’s benchmarks, frontier models still struggle with physics (e.g. 47.3% with GPT5.6-sol on Humanity Last Exam - Physics component). But that picture did not match physicists' experience using these models. So we did a rigorous analysis of current benchmarks and re-graded their supposed failures with Yale physicists. We found that the problem was often the benchmark and not the model !! When corrected, the models almost saturate all benchmarks, including everyone’s favourite, Humanity Last Exam (physics-part) and Critpt. These failures in turn make AA less reliable as a measure of model’s true capability. For more details, results and analysis check out our preprint: and blog post:
顯示更多
0
26
474
53
轉發到社區