The benchmarking landscape is evolving:
1. Domain-specific evals built by the companies that know the workflows best
2. Agent environments extended beyond a container into sandboxed infrastructure (data-eng-bench has two versions: local DuckDB and remote Snowflake)
3. Public evals released by companies to prove the effectiveness of their agent product on the workflows their customers care about
4. Benchmarks maintained as living software, with the community submitting trajectories to the leaderboard and proposing new tasks