Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
The journey building VulcanBench Safety v1 begins, and with it, my first private repo VulcanConduct.
This is phase zero of my AI safety eval suite build out, more to come, will share as I go as always 🖖