Benchmarking models across effort levels and harnesses on eval suites with real coding tasks. Because you don't need High or Max effort as much as you think.
A lot have people have been asking me why I donโt get early access to all the new models like the cool kids.
Easy answer.
All the influencers who get early access to models, sharing glowing posts about how much they loved using the models on launch day, often with some example where the models really shines.
At VulcanBench, I just share the data, good or bad, and let that be my guide.
As long as I continue on this path, I think getting early access isnโt going to happen for me.
But thatโs okay, Iโd rather just be able to share the good, the bad, and the ugly, on new models, even if my results come out a week or two after launch.
Live long and benchmark ๐