註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

wh
@nrehiew_
eng primarily, ml mostly, research previously
加入 October 2023
104 正在關注    18.5K 粉絲
Super cool work! This also explains why some benches have an "artificial ceiling". For example, on DeepSWE, Epoch found 23 false negatives out of 131 tasks. This means a ceiling of roughly 79.6%, which actually matches the current (almost) saturated high score of ~74%
顯示更多
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
顯示更多
0
8
405
14
轉發到社區