This is such a great trend. More of this. Evals that actually matter in the real world.
Introducing Supabase Evals.
Our benchmark for how well AI coding agents build with Supabase. We run agents like Claude Code, Codex, and Open Code against real tasks and score what they do.
顯示更多