註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Charlie Marsh
@charliermarsh
@OpenAI. Building Ruff, uv, ty, and other high-performance Python tools with the @astral_sh team.
加入 March 2009
981 正在關注    50.9K 粉絲
Lots of good examples in these reports. For example: in DeepSWE v1.1, the verifier discards the agent's changes to some test files. But the model isn't told this will happen! Which leads to all sorts of 'irrelevant' failures...
顯示更多
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
顯示更多