가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

Zhuokai Zhao
@zhuokaiz
AI Research Scientist @Meta. Building scalable intelligence. PhD @UChicagoCS.
가입 April 2024
383 팔로잉 중    5.3K 팬
For years the most popular coding benchmarks rank models on whether they can finish a fully-specified task on their own. Our new benchmark, SWE-Together, instead turns that one-shot test into an interactive session, and scores the agent both on how well it solves the problem and on how much steering it takes to get there. To measure that, we collect 11,260 recorded sessions, filter for those with genuine multi-turn feedback, real agent-authored edits, and verifiable outcomes, and rebuild the survivors into reproducible tasks, where each task is reconstructed in a sandbox with the repo pinned at its original commit and the user's first message as turn one. A reactive LLM user simulator then replays each session. It stays anchored to the original user's intent but speaks only when the agent's own trajectory calls for it (e.g., a clarification, a correction, a new requirement) instead of firing on a fixed schedule, so the corrections are something the agent draws out rather than a script we impose. SWE-Together measures two things: 1. Final correctness — whether the final repo (after the user interventions) does what the user actually asked. 2. User Correction — how much the user had to steer to get there. We also track Intent Coverage, a check that the simulator put the same underlying requests to every agent, so differences in correction reflect the agents and not an inconsistent simulator. As the existing one-shot scores saturate, how little an agent makes you intervene — and how well it ends up where you actually meant — can be the new signal for which model is worth using, and that's what we built SWE-Together to measure. Benchmark page:
더 보기