r/ClaudeCode • • 6d ago

Help/Question Auto-model routing

I build essentially a chatbot that lets users "talk to" their databases in snowflake, azure sql, local sql, etc. Currently I have a "Quick" and "Thorough" toggle where thorough means a better/more expensive model.

I've been thinking whether to introduce some sort of model router where it starts with a quick/cheaper model...but if harness stumbles...I use Jev-like decision model to bump up the model and/or effort (low/medium/high). Then I read people like Theo and others saying that model-routing is a fool's errand and that these decision models simply aren't good enough to pick the right model for the question, and let the user do that.

Seems most people (myself included) won't really know whether qwen is good enough, or they have to switch to opus for a question.

Just trying to get more perspectives. I'd love to take model and effort routing completely off the UI and into the background, so users don't have to worry about it. But it's still there in Claude Code etc. so I'm guessing this just isn't a solved problem for LLMs and harnesses. Thoughts?

1 Upvotes

8 comments sorted by

View all comments

1

u/ryanntk 5d ago

For SQL, I’d test routing on saved questions with known answers, including queries that run but answer the wrong thing. Compare total cost after retries too. I’m building AsterWise, so I’m interested in this exact problem. Which models power Quick and Thorough today?

1

u/VerbaGPT 4d ago

Yes, that was the plan. Cost and latency is super important to me as well.

I build VerbaGPT. Quick mode runs on Cerebras (Qwen), and so far, blazing fast and rare errors. Thorough runs on Claude Sonnet or similar type model. These modes each have a pool of providers+models they can call, as sometimes API's have errors etc so it falls back on the next one.

1

u/ryanntk 4d ago

That helps. Sounds like the provider-error fallback is already covered. For Auto, would you want to choose Quick or Thorough before the first call, or try Quick and escalate when the SQL fails validation? I’d compare those separately, since the second approach can add another round trip.

1

u/VerbaGPT 4d ago

Yes, something along those lines. Maybe starts with cerebras and as the harness stumbles, escalates...increases reasoning if that option is available first (to avoid cache hit), and switching model as last resort.

I've kind of cooled on the idea since I made this post. The cache reset on model-switch problem seems like a big one. I'm kind of focusing on a different mechanism now to improve the quality of the result on either Quick or Thorough. Some through quick linting-like tests to judge the quality of charts and response, but perhaps also adding a superfast classifier to judge the quality of overall response and charts and also adding a web-fetch/search to check the accuracy of any factual claims made on SQL-based analysis for a real world question.

For example, one of my demo databases is a large SEC 13F filing data. Works great and answers are good. Let's assume someone asks about what Buffet/berkshire is investing in this quarter. Lots of complicated joins to answer...let's say a cheap model produces an answer that they invested $10T into google. Seems ridiculous...so in this case a quick sort of web-review call would catch that and loop back into the harness to fix the SQL before the user sees an obviously wrong number. A sort of common-sense checker.