Register and share your invite link to earn from video plays and referrals.

Charlie Marsh
@charliermarsh
@OpenAI. Building Ruff, uv, ty, and other high-performance Python tools with the @astral_sh team.
Joined March 2009
981 Following    50.9K Followers
Lots of good examples in these reports. For example: in DeepSWE v1.1, the verifier discards the agent's changes to some test files. But the model isn't told this will happen! Which leads to all sorts of 'irrelevant' failures...
Show more
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
Show more