登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Tomas Hernando Kofman
@tomas_hk
building ¬◇, intelligent model routing for coding agents
参加 April 2012
794 フォロー中    3.4K ファン
We've been working on intelligent model routing since 2023 and partner with many of the F100 to reduce their coding agent costs. It is a deceptively difficult problem with many pitfalls that people tend to miss until they fall right into them. When new models become even just a little bit better, they are able to run much longer autonomously / in parallel. This means that subtle variations in the model landscape will lead to major changes in your inference spend, even if model prices don't go up. Model routing helps ensure you are not overpaying for the entire course of additional workload volume at the price of the bleeding frontier. But naive approaches, most notably complexity and semantic classifiers, will fail in agentic settings and generally cost you *more* money because they are not cache-aware and they dumbly look at just the next turn without considering all downstream impacts of model recommendations over the full horizon of the trajectory. On top of this, they tend to be heuristically constructed, basically shifting the burden of manual model selection away from the user and to the opaque opinions of the router designer who is far from the front line. At Not Diamond, we approach model routing for coding agents in a more data-driven way that accounts for the multiple interrelated variables that drive cost. By doing so, we are able to consistently deliver 20-30%+ savings on coding agent workloads with no degradation relative to the frontier. We've written more about our approach, along with the shape of the problem space more generally, here:
もっと見る