I think a lot about model routing, heck I even write a Substack just it, so yeah, I'm kinda obsessed.
At a high level, I don't think most of the model routers out there make much sense for anyone but enterprise co's already paying api pricing to use services like Anthropic and OpenAI.
And I think this is honestly, the target market for most of the companies releasing model routers today. So it makes sense, and yes, will definitely save enterprises money.
If you're spending $500,000/year on a Claude Sub. And everyone is just using Opus High for everything, yup, putting some thought into model selection, even at API+ pricing, which is what most routers charge, is still a savings.
But if you're using a subscription, at $100 or $200/mo, switching to a model router will dramatically increase what you pay, because you're not paying API pricing, you're paying a massively discounted price.
This doesn't mean the model routing co's are doing anything wrong, they're just not solving a problem for you, they're solving a problem that big enterprise co's have.
I've seen so many posts, and talked to so many people at big co's that say, it's all too overwhelming for them and their teams, so they tend to just see everyone using one model, usually the latest frontier model, and at High effort for everytihng.
What will be interesting to see is over time will these companies just decide to outsource this decision logic to third party model routers, or will they use benchmark data, both public, and internal evals, to just give their teams some model routing decisions, i.e. use these three models, at these effort levels, etc.
I think the biggest, unexplored territory right now is effort level. This is why I started
@VulcanBench because I saw just about every benchmark just running evals with models at Max effort, and I knew for myself and my team, we used Medium effort more than any other effort level.
But a model that does great at Max, might not be great at Medium, and another model might be better. And it could be completely different for Python code than Rust code, and medium-sized codebases vs. large.
These days, I wake up every morning so excited to be doing research in this area. There is so much to discover.
And that's my Monday brain dump, big week ahead, TGIM 🖖