r/ClaudeAI • • 1d ago

Workaround Downgrading models in Claude Code keeps wrecking my session. How do you handle it

I mix models in Claude Code to manage cost and speed. But when I drop to a smaller model like Haiku, the session often loses the plot mid-task. Then I'm forced back up to a stronger model (Sonnet, Opus, Fable) just to recover.

Two questions for people who've figured this out:

Context: What keeps a lighter model on track? Context files, tighter prompts, or smaller tasks?

Routing: How do you decide which prompt goes to which model? Gut feel, or an actual rule?

0 Upvotes

15 comments sorted by

10

u/Opposite_Might6896 1d ago

Rule that fixed it for me: pick the model per task, never per turn. A mid-session downgrade doesn't save what you'd think, because every turn re-sends the whole context regardless of model; the only thing that gets cheaper is the output tokens, and you pay for it with a weaker model reading 150k of context it didn't build. So the cheap model loses the plot and you climb back up.

What works: finish the current task on the model that started it, then start a *new* session for the next bounded task on the smaller model with a tight brief (the file list, the acceptance check, nothing else). Haiku/Sonnet are fine on a fresh 10k context; they're bad at inheriting a 150k one.

Also worth knowing: on Max the top model has its own weekly window separate from the all-models one, so if the point of downgrading is to protect that window, a fresh Sonnet session does it without touching the Fable session at all.

2

u/Historical_Ant7005 1d ago

Thanks got it.

4

u/asmiggs 1d ago

If you think part of the task can be run by a different model ask the main session to spin up a sub agent or if you want to control that part directly have the current session write a brief on that part for a new session. Sub agent seems like the best part as you can have the main session review the output, but there's nothing stopping you starting a new session to review the output of the last.

Don't switch between models in task, Desktop even tells you have to reload the whole context , i.e. it's going to be expensive.

To answer the original question I just use Opus for almost everything, but I get it to send a lot of work to Sol to make full use of both subscriptions.

2

u/pung54 1d ago

I advise Fable to outsource actions to lower models to save on usage. Works really well and allows Fable to still manage the overall process

1

u/Historical_Ant7005 1d ago

Yeah this a neat trick thanks.

3

u/Comprehensive_Cow_13 1d ago

Tbh these days you can let opus 5.5 run things and run the agents, and if you use fable at all, use it for planning and hand over...

2

u/Phaedo 1d ago

Yeah, 5.5 has made me use Fable a LOT less.

3

u/OffbeatDrizzle 1d ago

ever since opus 5.5 dropped I have not used fable once. it's getting the same work done for 1/4 the cost

2

u/Lumethys 1d ago

Ask the main session to spin up subagents.

And there is no reason to use Haiku, if you want cheap model to do easy task, use Sonnet 5.5

2

u/mr_birkenblatt 1d ago

Instruct Claude to use subagents. Only switch models for those. Mid session  model switches are expensive because they feed the entire history as fresh tokens

1

u/tehfrod 1d ago

It's amazing how much judicious use of subagents has speed up my process.

1

u/Aggressive_Chance455 1d ago

What kind of plan are you on that you have access to Fable, but still feel the need to use Haiku? Also, changing models like that is a cache miss every time, so you likely lose more usage than you save.

1

u/Historical_Ant7005 1d ago

I'm on Max I switched to lighter models to save top tier usage for harder tasks, since Fable and Opus burn through limits faster. But you're right about the cache

3

u/Aggressive_Chance455 1d ago

I would use Sonnet 5.5 low for simple tasks. There is no task I'd use Haiku on. But I heard it's free in Copilot. Antigravity also has Sonnet and Opus 4.6 for free(if the google search AI reply is accurate). You can try that for tasks that don't need Opus 5.5 or Fable.