One of the biggest flaws in coding benchmarks today is that real users interact with coding agents interactively, but benchmarks test only a single turn.
VCB 1->100, which we're releasing today, aims to change that.
Existing coding benchmarks stop at the first working version. We are releasing Vibe Code Bench 1-100 today, to measure what comes next.
This benchmark asks if models can handle a large number of modifications to a product, without breaking what already works.