가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Dimitris Papailiopoulos
@DimitrisPapail
Researcher @Microsoft | Prof @UWMadison (on leave) | babas of Inez Lily.
가입 May 2012
1.5K 팔로잉 중    29.8K 팬
ok i kinda love @terminalbench folks, so sorry for being a little COI'd here, but was very surprised to see this because my knee-jerk interpretation of the post was "TB 4.0 is broken". Epoch indeed did not say broken, they found 30/66 have scoring defects, but I feel the way this is presented it kinda sounds like "it's broken, don't trust it". Reading the review there do seem to be some real issues worth fixing eg exploitable graders, answer leakage and cases where correct solutions can be rejected. BUT "30/66 tasks have scoring issues" is not the same as "45% of TB4 results are wrong." Grader being exploitable doesn't tell how OFTEN it was exploited (yet indeed this needs fixing) and what it means for the leaderboard. I guess the takeaway is that there are defects that should be taken seriously but one one should be very careful with wording because "this eval has flaws" can be read as yet IS NOT the same as "this benchmark is broken don't trust it." (this is not a dunk on Epoch, they are doing great work !)
더 보기