注册并分享邀请链接,可获得视频播放与邀请奖励。

Kevin Li
@kevin_x_li
Building Harbor-Index, Prev: Stanford, Nvidia, UMich
加入 March 2022
300 正在关注    367 粉丝
"we chased after benchmarks, when none of the benchmarks measure whether humans actually enjoy working with the model" We actually DO have benchmarks for the user experience of coding agents! Check out SWE-Together by @yifannnwu et al. It replays real user-agent interactions and measures not just pass rate but user effort: the number of corrective feedback turns needed to keep the agent on track.
显示更多