I have now had more than a few companies reach out to me about doing stack-specific benchmarks and model reports with
@VulcanBench.
This wasn't something I thought of, but then it kinda clicked.
I'm also hearing now more and more about challenges eng leaders are having taking all the benchmarks that are out there, and figuring out how to apply them to their teams, and what kind of model routing might be best for them based on their stack and codebase size.
This is also where effort levels really matter. I've heard from a lot of people that, "my team just uses one model, high effort, for everything."
And yeah, that's not optimal. It's both expensive, and slow, and you'll end up with your eng team waiting around for hours, when they could get a task done faster, and cheaper, if they had better model routing.
This is different from all the model routers companies are coming out with, because those are general, with VulcanBench, I can get really specific using actual PRs from real work that eng teams are doing to customize model routing just for them.
And, since I now have a growing list of companies that want this, I guess it just makes sense to add it to the VulcanBench site and start offering it as a service, and we'll just see where it goes!
No idea how to price it, just putting it there as an email link and will take it convo by convo.
You can read more about it here: