We created data-eng-bench with
@bespokelabsai and we’re open-sourcing it.
There are plenty of model performance benchmarks for code generation. But for data engineering, the harness matters just as much as the model.
We built a benchmark that asks agents to build and fix real pipelines, then grades them on whether the output actually works.
The results: Using the same Opus 5 model,
@Snowflake CoCo achieves 73.8% Pass
@1 at 3.9× lower cost than Claude Code. With Sonnet 5, CoCo delivers the same quality at 2.3x lower cost.
Now anyone can use data-eng-bench to evaluate model + harness combinations on real data engineering workflows. Run your own tests and share what you find.