Register and share your invite link to earn from video plays and referrals.

Search results for ModelRouting
ModelRouting community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including ModelRouting
TL;DR: No single LLM is optimal for every query and budget. This work unifies the scattered field of LLM routing into "five building blocks" and ships an open-source infrastructure bundling 16+ routers with a dedicated benchmark. Title: LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers URL: Key points 🧩 Decomposes any router into 5 components: Context Encoder / Model Encoder / Scoring / Decision Rule / Learning Signal 📊 xRouteBench spans 5 tracks (Generic, Memory, Vision, TimeSeries, Personalized), 4,767 queries total 🛠️ A dense query-model matrix, built by dispatching every query to all 18 candidate models, serves as both training supervision and test bed 🚀 Learned routers beat the strongest fixed-model baseline (always pick the largest) by 14.6% relative ⚖️ No router dominates: RouterDC leads on accuracy but drops to 10th under tight cost constraints 🔁 Multi-turn routing adds cost without consistent gains (Router-R1 at just 22.3%) 👤 User conditioning helps, but the best design flips between simulated personas and real feedback A solid step that proves "the best router depends on task and budget" and lays a foundation for fair comparison. #LLM# #ModelRouting#
Show more
Introducing model routing to Factory. Factory Router picks the right model for every task, automatically. Maintain frontier performance while cutting costs by 25%.
0
198
2.2K
210
Forward to community
This new Nvidia paper is huge for model routing and the future of inference: re-using the KV cache across different LLMs. One of the greatest challenge in model routing during long-horizon agent settings is that the cache does not transfer across different models. As a result, naively switching models without KV-cache-awareness will result in wasting more money. Conversely, transferring the KV cache across models significantly expands the total amount of savings that can be achieved. This paper is a great step in that direction. Congrats to Taekyung Heo, Rasoul Shafipour, Ritchie Zhao, Maximilian Golub, @MMKamani7, @Ritika_Borkar, Makesh Tarun Chandran, Pantea Zardoshti, and Bita Darvish Rouhani. Two questions: 1. Do you plan to investigate cross-family transfer? I saw this in the future work section but curious to learn more about your intuitions on the opportunities and challenges here. 2. What was the reason for focusing mostly on small to large model transfers and not large to small?
Show more
I'm giving a talk on model routing today at #Ai42026#. If managing AI inference spend is a priority for your organization, please drop by! I'll also be available after the talk for Q&A. @Ai4Conferences
Show more
We've been working on intelligent model routing since 2023 and partner with many of the F100 to reduce their coding agent costs. It is a deceptively difficult problem with many pitfalls that people tend to miss until they fall right into them. When new models become even just a little bit better, they are able to run much longer autonomously / in parallel. This means that subtle variations in the model landscape will lead to major changes in your inference spend, even if model prices don't go up. Model routing helps ensure you are not overpaying for the entire course of additional workload volume at the price of the bleeding frontier. But naive approaches, most notably complexity and semantic classifiers, will fail in agentic settings and generally cost you *more* money because they are not cache-aware and they dumbly look at just the next turn without considering all downstream impacts of model recommendations over the full horizon of the trajectory. On top of this, they tend to be heuristically constructed, basically shifting the burden of manual model selection away from the user and to the opaque opinions of the router designer who is far from the front line. At Not Diamond, we approach model routing for coding agents in a more data-driven way that accounts for the multiple interrelated variables that drive cost. By doing so, we are able to consistently deliver 20-30%+ savings on coding agent workloads with no degradation relative to the frontier. We've written more about our approach, along with the shape of the problem space more generally, here:
Show more
xAPI now supports URL-level model routing. Point any OpenAI or Anthropic-compatible request — Claude Code or Codex — to: and xAPI will automatically translate the protocol and route the call to DeepSeek V4 Pro (you can replace it with any other models). One endpoint. Any standard. Fixed target model:
Show more
499 Detalks [CN] 225: Deconstructing the Agentic Web: Multi-Model Intelligent Routing & On-Chain Financial Infrastructure Accelerating AGI Landing Time:28, April, 2026, 20:00 (UTC+8) Hosted by 499 and co-hosted by @BAI_AGI, this Space explores how intelligent model routing, on-chain data infrastructure, and dedicated agent financial rails can bridge the gap and accelerate AGI into the real world. Mod: Charis @charis_em, 499 Core Member Guest Speakers: Jtsong @Jtsong2, Head of APAC, @0G_labs Anita @Anitahityou, APAC, @SentientAGI Leslie @leslieloser_ , Content Creator Mia @Artistkatty_ , Growth from Starchild Rika @rayrayweb5, KOL Key discussion topics include: -- From “chat tools” to “autonomous economic entities”: What is the single biggest bottleneck stopping AI Agents from true mass adoption -- Multi-model intelligent routing & data orchestration: How far has the industry come in routing across multiple models and data sources -- Agent Wallets & on-chain credit systems: The power of Machine-to-Machine (M2M) value transfer when every AI has its own wallet Space Link: Telegram Group:
Show more
I have now had more than a few companies reach out to me about doing stack-specific benchmarks and model reports with @VulcanBench. This wasn't something I thought of, but then it kinda clicked. I'm also hearing now more and more about challenges eng leaders are having taking all the benchmarks that are out there, and figuring out how to apply them to their teams, and what kind of model routing might be best for them based on their stack and codebase size. This is also where effort levels really matter. I've heard from a lot of people that, "my team just uses one model, high effort, for everything." And yeah, that's not optimal. It's both expensive, and slow, and you'll end up with your eng team waiting around for hours, when they could get a task done faster, and cheaper, if they had better model routing. This is different from all the model routers companies are coming out with, because those are general, with VulcanBench, I can get really specific using actual PRs from real work that eng teams are doing to customize model routing just for them. And, since I now have a growing list of companies that want this, I guess it just makes sense to add it to the VulcanBench site and start offering it as a service, and we'll just see where it goes! No idea how to price it, just putting it there as an email link and will take it convo by convo. You can read more about it here:
Show more
🦞 OpenClaw 2026.6.10 just dropped. Just a small release to keep things brewing: ⚡ Automatic fast mode for short talks 🧠 Much more reliable model routing 🔒 Safer session state + trusted policies 🛠️ Better provider onboarding Helping deliver rock-solid lobsters. 🦞
Show more
AiMaMi: OpenAI Codex Desktop Manager Organize sessions, MCP configs & model routing. 👉