註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Yafah Edelman
@YafahEdelman
Chief Strategy Officer @EpochAIResearch she/her
加入 March 2014
618 正在關注    2.4K 粉絲
We audited 15 benchmarks and labeled 9 flawed: - In Terminal Bench 4.0 we found 45.5% of tasks to be broken after reviewing github issues. - In HLE, 46% of the 48 questions we randomly sampled were broken. - In DeepSWE 1.1 we found a bug that can break grading for every task.
顯示更多
0
42
640
41
轉發到社區