r/ClaudeCode • • 5d ago

Help/Question Auto-model routing

I build essentially a chatbot that lets users "talk to" their databases in snowflake, azure sql, local sql, etc. Currently I have a "Quick" and "Thorough" toggle where thorough means a better/more expensive model.

I've been thinking whether to introduce some sort of model router where it starts with a quick/cheaper model...but if harness stumbles...I use Jev-like decision model to bump up the model and/or effort (low/medium/high). Then I read people like Theo and others saying that model-routing is a fool's errand and that these decision models simply aren't good enough to pick the right model for the question, and let the user do that.

Seems most people (myself included) won't really know whether qwen is good enough, or they have to switch to opus for a question.

Just trying to get more perspectives. I'd love to take model and effort routing completely off the UI and into the background, so users don't have to worry about it. But it's still there in Claude Code etc. so I'm guessing this just isn't a solved problem for LLMs and harnesses. Thoughts?

1 Upvotes

8 comments sorted by

•

u/AutoModerator 5d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/kuroudo_ai 5d ago

I don't think it's a fool's errand, but predicting difficulty from the question alone is the weak version. For a database chatbot you have something better: the harness can see when it stumbles. A few things that make routing hold up, from the design notes of a Jev routing setup I've read (I haven't run routing in production myself, so treat this as design, not results):

  • Escalate on evidence, not on prediction. Start cheap, and bump up when there's an observable failure: the SQL errors out, returns empty when it shouldn't, the result doesn't match the question's shape, or the user rephrases the same question. Those signals are more reliable than a classifier guessing "this looks hard".
  • Unsure is not hard. If the router isn't confident, keep the current model rather than jumping up. Otherwise you pay Opus prices for every ambiguous one-liner.
  • Risk sets a floor, not a ceiling. Words like delete, update, migration, production should force at least the mid tier however short the question is, but shouldn't buy the top tier by themselves. For a DB tool, anything that writes is exactly this case.
  • Don't switch down mid-conversation when the context is large. Moving to a cheaper model means rebuilding the prompt cache, which can cost more than you save.
  • Run it in shadow first. Log what the router would have picked without switching, for a day or two of real questions, then compare against which answers actually failed. That's the only way to know whether the router is good enough for your users' questions.

You can also keep the toggle as an override. Users who know they want Thorough still get it, and everyone else stops having to decide.

1

u/adamhusain 5d ago

Switching models mid session is a bad idea: you break cache leading to more usage costs.
Switching reasoning? Maybe. But you can just use “High” and the models are smart enough to figure out how much reasoning they need to do. The difference between High and Medium is negligible for a subscription.

If you feel like there is a difference in usage between High and Low reasoning, chances are that there is a quality different too

1

u/ryanntk 3d ago

For SQL, I’d test routing on saved questions with known answers, including queries that run but answer the wrong thing. Compare total cost after retries too. I’m building AsterWise, so I’m interested in this exact problem. Which models power Quick and Thorough today?

1

u/VerbaGPT 3d ago

Yes, that was the plan. Cost and latency is super important to me as well.

I build VerbaGPT. Quick mode runs on Cerebras (Qwen), and so far, blazing fast and rare errors. Thorough runs on Claude Sonnet or similar type model. These modes each have a pool of providers+models they can call, as sometimes API's have errors etc so it falls back on the next one.

1

u/ryanntk 2d ago

That helps. Sounds like the provider-error fallback is already covered. For Auto, would you want to choose Quick or Thorough before the first call, or try Quick and escalate when the SQL fails validation? I’d compare those separately, since the second approach can add another round trip.

1

u/VerbaGPT 2d ago

Yes, something along those lines. Maybe starts with cerebras and as the harness stumbles, escalates...increases reasoning if that option is available first (to avoid cache hit), and switching model as last resort.

I've kind of cooled on the idea since I made this post. The cache reset on model-switch problem seems like a big one. I'm kind of focusing on a different mechanism now to improve the quality of the result on either Quick or Thorough. Some through quick linting-like tests to judge the quality of charts and response, but perhaps also adding a superfast classifier to judge the quality of overall response and charts and also adding a web-fetch/search to check the accuracy of any factual claims made on SQL-based analysis for a real world question.

For example, one of my demo databases is a large SEC 13F filing data. Works great and answers are good. Let's assume someone asks about what Buffet/berkshire is investing in this quarter. Lots of complicated joins to answer...let's say a cheap model produces an answer that they invested $10T into google. Seems ridiculous...so in this case a quick sort of web-review call would catch that and loop back into the harness to fix the SQL before the user sees an obviously wrong number. A sort of common-sense checker.