Introducing Vals-Smith: turn your code base into a customized benchmark.
Public benchmarks tell you which model is strongest overall, not which model is the best on your code. Vals-Smith turns your merged pull requests into real coding tasks and measures the percentage a model can actually resolve.
New models ship every week. Vals-Smith tells you which one to trust with your code.