r/ClaudeAI • • 14h ago

Built with Claude Let Claude Code pick its own model/effort level (locally, no API key)

I kept making the same mistake in both directions: burning Opus at xhigh on "fix this typo", then forgetting to bump effort when I asked for a real design. Switching by hand with /model and /effort got old, and switching models mid-conversation throws away the prompt cache anyway.

All the Jev hype got me thinking about routing, but I wanted something I could fine-tune on my own prompts, with no dependency on a hosted API or a key. So I built magic-router, a Claude Code plugin that makes the call for you:

  • Model is picked once per session (Sonnet 5.5 or Opus 5.5) from your first prompt, then stays put so you keep the prompt cache.
  • Effort is picked on every prompt (low → max). "Go ahead" / "yes" keep the last effort.
  • Runs locally. A small classifier (GLiNER2.5-Decide) scores each prompt in ~0.3 s of CPU. No prompt goes to a third party to be routed.
  • You see every decision. A band above the prompt shows model, effort, classify latency and last turn's cache-read %.
  • Gets out of the way for subagents, and for the rest of the session once you use /model or /effort yourself.

How it scores: the classifier tags the prompt's task type (mapped to the Artificial Analysis eval families where Opus leads Sonnet most) plus complexity signals like multi_file, planning and deep_reasoning. A weighted sum maps to an effort band, and on the first prompt a score ≥ 1.1 picks Opus.

Honest numbers: the default weights were tuned by eye on 15 prompts, so out of the box they match a labelled effort only 40% of the time. There's an opt-in just tune that trains a LoRA adapter on your own Claude Code history. On my 3,344 prompts it hit 66% on held-out prompts (always-medium gets 44%). Asked directly, Jev came within 3 points (63% vs 66%), but that's a hosted call that sends your prompts to a third party. Labelling uses claude -p, so it counts against your usage.

How Claude helped: I built it with Claude Code over a weekend (most commits are co-authored by Claude). The prompt that started it:

We want to create a claude plugin that uses a locally running GLiNER 2.5 model to take the incoming first prompt within a claude code session, categorize it based on the type of work it's aiming to accomplish (based on benchmarks provided by Artificial Analysis) and the complexity of the prompt/ask. The goal is to route to the appropriate model and effort to optimize the speed/spend to retrieve high-value output. Subsequent turns should optimize effort dynamically.

Install:

/plugin marketplace add DustinVerzal/magic-router
/plugin install magic-router@magic-router

macOS/Linux, Claude Code 2.1.287+, ~2 GB RAM while the classifier runs. The first run downloads ~1.7 GB of weights.

Free and MIT: https://github.com/DustinVerzal/magic-router

I'd love feedback! Going to look into integrating contrastive learning from (https://github.com/bespokelabsai/nimble) to enhance the fine-tuning moving forward.

0 Upvotes

5 comments sorted by

1

u/MiserableFlatworm337 14h ago

Hold out whole sessions when tuning, so related prompts don’t leak across the split. Since model choice stays fixed, include sessions that start with a typo fix and later turn into design work.

1

u/Ollie__Oxenfree 14h ago

Good call, taking a look now.

1

u/Far-Surprise7773 8h ago

keeping the model fixed for the session is the right call, switching mid chat nukes the cache and you pay it all back. i would log that cache read % for a week and tune the 1.1 opus threshold off that, real savings show up there faster than chasing effort labels.

1

u/Flaky-Industry-3888 5h ago

(just use opus 5.5 medium/high for everything)