Been looking into token optimization and model routing, I think super obvious optimization to tackle both cost + demand on inference
Here’s a small post about different techniques and methods
Introducing model routing to Factory.
Factory Router picks the right model for every task, automatically.
Maintain frontier performance while cutting costs by 25%.