Okay, my new 23-task eval suite for
@VulcanBench is done. A lot of little details I wanted to get right with this.
It was important to me that I had more broad language coverage, and also that it takes into account the issues that OpenAI raised yesterday with SWE-Bench Pro's four failure modes.
I'm new to benchmarking, but learning more every time I create a new set of evals. I'm calling this set v3, getting ready to run the first smoke tests.
Then this weekend, phew, safe to say I have a lot of models to test!
I will be committing these to Github so anyone can take a look and give me feedback on these as well.
Live long and Benchmark 🖖