登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Yifan Wu
@yifannnwu
吴奕凡; AI Research Scientist @Meta | Ph.D. @penn @picslupenn @GRASPlab.
参加 November 2016
490 フォロー中    1.4K ファン
This matches exactly what we saw building SWE-Together. For real-world coding tasks, a fixed test suite rarely captures the whole truth. Generated tests can be too strict or low-coverage, too tied to the reference patch, or too shallow about whether the feature actually works. And a failing run can come from environment noise rather than a true model capability gap. So we added an agentic judge: we freeze a weighted behavioral rubric from the task spec, user intent, and oracle trajectory, then run it in a fresh task sandbox to inspect the patch, workspace, and tests against those goals. Separating model capability from eval noise takes careful design and audits. The next hard problem is maintenance: benchmarks seem to saturate every few months. We need better ways to refresh tasks reliably and automatically. How should the field approach this? See more at
もっと見る
We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval.
もっと見る