注册并分享邀请链接,可获得视频播放与邀请奖励。

Yafah Edelman
@YafahEdelman
Chief Strategy Officer @EpochAIResearch she/her
加入 March 2014
618 正在关注    2.4K 粉丝
We audited 15 benchmarks and labeled 9 flawed: - In Terminal Bench 4.0 we found 45.5% of tasks to be broken after reviewing github issues. - In HLE, 46% of the 48 questions we randomly sampled were broken. - In DeepSWE 1.1 we found a bug that can break grading for every task.
显示更多
0
42
640
41
转发到社区