Register and share your invite link to earn from video plays and referrals.

sridhar
@RamaswmySridhar
CEO @snowflake; founder @neeva Ex-@GreylockVC Ex-@Google SVP of Ads Ex-@BellLabs.
626 Following    35.8K Followers
We created data-eng-bench with @bespokelabsai and we’re open-sourcing it. There are plenty of model performance benchmarks for code generation. But for data engineering, the harness matters just as much as the model. We built a benchmark that asks agents to build and fix real pipelines, then grades them on whether the output actually works. The results: Using the same Opus 5 model, @Snowflake CoCo achieves 73.8% Pass@1 at 3.9× lower cost than Claude Code. With Sonnet 5, CoCo delivers the same quality at 2.3x lower cost. Now anyone can use data-eng-bench to evaluate model + harness combinations on real data engineering workflows. Run your own tests and share what you find.
Show more
But here's the punchline. Normalized to 90% cache hit rate: GLM-5.2 (Fireworks): $1.12/session Opus-4.7 (Anthropic): $2.14/session GLM is ~48% cheaper.
Follow-up to my GLM vs Opus thread: let's talk cost. We ran 103 dbt tasks x 3 trials on each model. Same harness, same tasks. GLM: 860M tokens Opus: 439M tokens That's ~2x. But the "why" is more interesting than the number.
Show more