注册并分享邀请链接,可获得视频播放与邀请奖励。

Charlie Marsh
@charliermarsh
@OpenAI. Building Ruff, uv, ty, and other high-performance Python tools with the @astral_sh team.
加入 March 2009
981 正在关注    50.9K 粉丝
Lots of good examples in these reports. For example: in DeepSWE v1.1, the verifier discards the agent's changes to some test files. But the model isn't told this will happen! Which leads to all sorts of 'irrelevant' failures...
显示更多
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
显示更多