Register and share your invite link to earn from video plays and referrals.

Yifan Wu
@yifannnwu
吴奕凡; AI Research Scientist @Meta | Ph.D. @penn @picslupenn @GRASPlab.
Joined November 2016
490 Following    1.4K Followers
This matches exactly what we saw building SWE-Together. For real-world coding tasks, a fixed test suite rarely captures the whole truth. Generated tests can be too strict or low-coverage, too tied to the reference patch, or too shallow about whether the feature actually works. And a failing run can come from environment noise rather than a true model capability gap. So we added an agentic judge: we freeze a weighted behavioral rubric from the task spec, user intent, and oracle trajectory, then run it in a fresh task sandbox to inspect the patch, workspace, and tests against those goals. Separating model capability from eval noise takes careful design and audits. The next hard problem is maintenance: benchmarks seem to saturate every few months. We need better ways to refresh tasks reliably and automatically. How should the field approach this? See more at
Show more
We audited SWE-Bench Pro, one of the most widely used AI coding benchmarks, and found it no longer reliably measures frontier coding capability. We find 30% of SWE-Bench Pro tasks to be broken, and are retracting our previous recommendation that the research community use it as a leading coding eval.
Show more