"we chased after benchmarks, when none of the benchmarks measure whether humans actually enjoy working with the model"
We actually DO have benchmarks for the user experience of coding agents! Check out SWE-Together by
@yifannnwu et al. It replays real user-agent interactions and measures not just pass rate but user effort: the number of corrective feedback turns needed to keep the agent on track.