註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Kevin Li
@kevin_x_li
Building Harbor-Index, Prev: Stanford, Nvidia, UMich
加入 March 2022
300 正在關注    367 粉絲
"we chased after benchmarks, when none of the benchmarks measure whether humans actually enjoy working with the model" We actually DO have benchmarks for the user experience of coding agents! Check out SWE-Together by @yifannnwu et al. It replays real user-agent interactions and measures not just pass rate but user effort: the number of corrective feedback turns needed to keep the agent on track.
顯示更多