AI is getting more capable and harder to measure.
Benchmarks test contrived tasks. Test sets leak. Leaderboards get optimized against. Labs grade their own exams.
@ValsAI builds benchmarks the way work is actually done: evaluations authored alongside practitioners on real professional tasks, with private test sets that can’t leak into training data.
As models, harnesses, and applications multiply, every serious adoption decision routes through the same question: what can these systems actually do? Vals is building the trusted answer.
We led the seed and are proud to continue partnering with
@RayanKrishnan,
@langstonnashold and the team as
@a16z leads their Series A.
Our investment thesis: