註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Yifan Wu
@yifannnwu
吴奕凡; AI Research Scientist @Meta | Ph.D. @penn @picslupenn @GRASPlab.
加入 November 2016
490 正在關注    1.4K 粉絲
This matches exactly what we saw building SWE-Together. For real-world coding tasks, a fixed test suite rarely captures the whole truth. Generated tests can be too strict or low-coverage, too tied to the reference patch, or too shallow about whether the feature actually works. And a failing run can come from environment noise rather than a true model capability gap. So we added an agentic judge: we freeze a weighted behavioral rubric from the task spec, user intent, and oracle trajectory, then run it in a fresh task sandbox to inspect the patch, workspace, and tests against those goals. Separating model capability from eval noise takes careful design and audits. The next hard problem is maintenance: benchmarks seem to saturate every few months. We need better ways to refresh tasks reliably and automatically. How should the field approach this? See more at
顯示更多
We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval.
顯示更多