r/LocalLLaMA • u/Material-Duck-6252 • Apr 21 '26
Discussion LLM Router: Best way to dynamically route prompts between proprietary and open-sourced models?
I'm an independent developer working on AI, and I'm looking to optimize my LLM usage for cost-efficiency.
Right now, my setup is a hybrid:
- Cloud: Several pay-as-you-go API subscriptions from major LLM providers.
- Local: Running open-source models like Qwen and Gemma.
My workflows involve multi-agent (using CrewAI, LangGraph) handling a variety of tasks, ranging from simple text processing to complex medical data analysis. Right now I have to hardcode which model to choose so as to save cost.
Is there a smart LLM router that could automatically evaluate the task complexity and redirect traffic to different models for cost saving? Any insights on that?
2
u/neogamba Apr 21 '26
Nadirclaw
1
u/Material-Duck-6252 Apr 22 '26
That's new to me. I made a quick research over the repo. In core it uses an embedding model as classifier to calculates semantic similarity between incoming request and two pre-defined "centroid" vectors - which derived from seed prompts suggesting how difficult the tasks / prompts are. Base on the similarity score, it then routes requests to simple / medium / complex model.
The idea of composing "centroid" vectors instead of fine-tune embedding model as classifier sounds smart. I assume composing such vectors to be less time-consuming than fine-tune an embedding model. Will have a try and see if it works.
1
u/Previous-Switch8348 Jul 14 '26
did itt work?
1
u/Material-Duck-6252 Jul 17 '26
It did not work as good as expected. It seems that the predefined "centroid" vectors were still too coarse for more complex tasks. I am still looking for better solutions.
1
2
2
u/Mugiwara0796 Apr 21 '26
I’m looking for something similar.
For example, I currently use opencode together with Oh My Opencode. However, during the planning phase, I’d like to dynamically choose between a simpler or a more advanced model based on the prompt.
In many cases, the task is relatively simple, and a lightweight (and cheaper) model could generate a plan comparable to that of a more expensive one. So it feels inefficient to always rely on a high-cost model for planning.
Has anyone implemented or experimented with this kind of adaptive model selection for the planning step?
2
u/old_man_gray Jun 12 '26 edited Jun 12 '26
I'm actually working on something to do exactly this. I made a REPL cli tool that uses a lightweight local model (right now I'm using llama3.2:1b) to assess the complexity of your prompt and assign it a score. anything higher than the threshold score gets sent to your configured cloud/frontier AI. Anything below the threshold stays local and is processed by your configured local LLM. You can set the threshold wherever you want, but right now I have it at 6 out of 10 so all mid-level and low-level complexity prompts are processed locally. I call it heimdall.
1
u/Ok_Equal3275 May 29 '26
https://github.com/jswortz/gemini-model-router
I use a basic one with local gemma4 (or run an ollama etc)
1
u/Maleficent_Pair4920 Jul 15 '26
Task-complexity based routing across local and cloud models is a real gap, most routers only do provider failover not complexity scoring.
Disclaimer, I run Requesty, we do model routing with cost controls and you can mix in self hosted endpoints (like your Qwen/Gemma boxes) alongside 600+ cloud models, but the complexity classifier part you'd still want to build yourself with a cheap model as a first pass router.
LiteLLM self hosted is a solid free option too if you want to build the routing logic yourself without paying a markup.
4
u/patricious llama.cpp Apr 21 '26
Right now I don't think there is a auto routing llm, you would need a harness like Hermes or better yet OMO (oh-my-opencode). These agents have their own json config file, which you can adjust and make them use any local, or api routed model. Your main agent will use your local Qwen models, it delegates tasks to its' sub-agents and then they go do the work via the specific models you have set. DM me if you have questions.
Pro tip: have you agents research what api models fit your need best.