Register and share your invite link to earn from video plays and referrals.

Kevin Li
@kevin_x_li
Building Harbor-Index, Prev: Stanford, Nvidia, UMich
300 Following    367 Followers
"we chased after benchmarks, when none of the benchmarks measure whether humans actually enjoy working with the model" We actually DO have benchmarks for the user experience of coding agents! Check out SWE-Together by @yifannnwu et al. It replays real user-agent interactions and measures not just pass rate but user effort: the number of corrective feedback turns needed to keep the agent on track.
Show more
Introducing SWE-ZERO-12M-trajectories: the largest agentic trace dataset in the open, 5.7x larger than the previous largest. 112B tokens · 12M trajectories · 122K PRs · 3K repos · 16 languages
Show more