註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

sridhar
@RamaswmySridhar
CEO @snowflake; founder @neeva Ex-@GreylockVC Ex-@Google SVP of Ads Ex-@BellLabs.
加入 October 2018
626 正在關注    35.8K 粉絲
We created data-eng-bench with @bespokelabsai and we’re open-sourcing it. There are plenty of model performance benchmarks for code generation. But for data engineering, the harness matters just as much as the model. We built a benchmark that asks agents to build and fix real pipelines, then grades them on whether the output actually works. The results: Using the same Opus 5 model, @Snowflake CoCo achieves 73.8% Pass@1 at 3.9× lower cost than Claude Code. With Sonnet 5, CoCo delivers the same quality at 2.3x lower cost. Now anyone can use data-eng-bench to evaluate model + harness combinations on real data engineering workflows. Run your own tests and share what you find.
顯示更多
0
7
112
17
轉發到社區