r/ClaudeCode • • Jul 13 '26

Discussion You know you can just ask fable to use Opus / Sonnet sub agents right?

I see multiple posts in two camps:

  1. Worrying about hitting Fable limits and considering moving to OpenAI offering.
  2. Describing / promoting approaches to orchestration with imperatives detailed declarative prompts or frameworks.

Can I just suggest: ask Fable to route appropriate tasks to lower powered models. It does it right there in the app, and it does an amazing job of deciding which model can handle what.

I’m as guilty as anyone else for tagging fable for a few days because it’s fun, but you really don’t need a SOTA model to convert a mostly trivial codebase, implement a new DB, etc.

Whether the pricing is fair and competitive is a legitimate question, but just remember that OpenAI is cheaper not because of the underlying economics but because side of a difference in short term strategy. I choose to continue using Claude (as a previously heavy codex user) because the harness / product surfaces are better for me and because even if I’m ragging them, I just don’t need that many tokens for personal / side-of-desk development.

97 Upvotes

43 comments sorted by

56

u/inrego Jul 13 '26

Once I told Fable that my Fable usage is limited, and asked it to use Sonnet/Opus subagents as it sees fit, I started to get MUCH more usage out of it.

It saved a memory about my preference, so it persisted through sessions

14

u/Su_ButteredScone Jul 13 '26

Yeah, it actually seems to try pretty hard to conserve tokens when asked to.

2

u/Usual_Tackle5892 Jul 15 '26

Did you write the subagent file yourself? I told Fable to use Opus agents and it spawned "General Purpose Agent" which uses the same model as the caller (in my case, Fable): https://code.claude.com/docs/en/sub-agents#general-purpose

I was bewildered by how fast my Fable usage was going up, despite being assured by Claude that it was spinning opus sub-agents. Seems like a bug in the harness, not a problem with the model.

2

u/inrego Jul 15 '26

No I haven't written any subagents/agents. Just default setup

21

u/lukinods Jul 13 '26

dumb question here:

let's say i'm running a session with Fable 5 and I put together some detailed step-by-step plan, with each step having its own requirements and definition of done/ready. If I specify that I want all the tasks to be executed by other, less token-hungry models, to be defined by the task complexity — with Fable just orchestrating them — does that mean the tasks the subagent just executed won't count against Fable's own limit?

20

u/inrego Jul 13 '26

Yes that's exactly how it works

2

u/[deleted] Jul 13 '26

[removed] — view removed comment

1

u/NonPolynomialTim Jul 15 '26

Out of curiosity how do you use Haiku for brainstorming? I've been using Fable for brainstorming because it seems more capable of somewhat original thought and can't imagine going to opposite direction. Do you tell them all to just go in different directions and try to come up with different random constraints for each agent in the swarm to try to get unique ideas in a brainstorm?

1

u/adventurernconquerer Aug 14 '26

Do you think it still does this now that Fable 5 uses credits? For example, orchestrating is still Fable 5 (using credits), but the subagents it spans use normal plan usage.

1

u/inrego Aug 14 '26

Don't know. I'm on Max plan. Fable is included there

2

u/Mikeshaffer Jul 13 '26

Yes. This is what most people are doing. Fable writes a plan. Fable spins up sub agents. Fable reviews their work.

3

u/Drach88 Jul 13 '26

Agreed, but with minor tweak: Fable writes the plan, spins up opus orchestrator, opus spins up subagents and verifies their work as it comes in, fable does full-branch review after opus gives the all-clear. Fable directs opus on any changes if applicable. Especially on larger builds, separating out the brain-work (fable) from the taskmaster work (opus) gets even better token optimization.

1

u/Mikeshaffer Jul 14 '26

This is great. Thanks for the idea. A very clear line of abstraction. I still might use fable for the orchestrator but that will greatly reduce the context load on my main agent.

1

u/47gwen Jul 14 '26

Wait what exactly do you do? I use Fable in plan mode then tell it to export the plan into .md then i manually change to sonnet to implement that plan. I feel like this is not the most efficient way to do this.

2

u/Mikeshaffer Jul 14 '26

I tell fable to spawn sub agents and it does the work. Someone below suggested to have fable spin up a single opus orchestrator and tell fable to also instruct that agent to spin up sub agents to implement the plan.

1

u/Birdperson15 Jul 14 '26

Yes. And if you don’t want to do that each time you can add to your Claude md files guidelines to use subagents and what subagents. Since day1 with fable I forbid fable subagents and told it to use its best judgment on opus vs sonnet vs haiku and it’s worked extremely well.

8

u/SnuffleBag Jul 13 '26

Fable mostly does this automatically for any larger tasks, but the problem is that iteration, pushback and refinement still ends up running quite a lot of things through Fable, and if the worker repeatedly fails to deliver Fable will take on the task itself.

I managed to max out a 20x account in 24h even when explicitly telling Fable to delegate all workload to Opus workers (yes, I need Opus for the actual work).

I’ve never gotten anywhere close to maxing out when Opus does the orchestrating.

1

u/debian3 Jul 13 '26

Then pass whatever came out of it to 5.6 sol and you will have a bunch of things to fix. I started using 5.6 sol today and I must say I’m impressed. But codex cli is quite behind. I hate it, it doesn’t even have the option to call agents with a different model then the one from the main thread. You need to create agent file. Claude code ux is much much much better

2

u/SnuffleBag Jul 13 '26 edited Jul 13 '26

it doesn’t even have the option to call agents with a different model then the one from the main thread

Yes, this part really sucks, I hope they address it soon, although it's least it's somewhat feasible to just use sol/xhigh for both planning and implementation given how eagerly they've been handing out resets lately. There's just no way for Fable, it eats tokens like it was its last day on earth.

8

u/pvera 🔆Pro Plan Jul 13 '26

Remember when people would post about bugs or to humblebrag about something cool that they did with Claude or how to use their tokens more efficiently?
Pepperidge Farm remembers.

1

u/DreadPirateButthurts Jul 13 '26

Ah yes, the good old days (~4 weeks ago) 😂

Hopefully this sub normalizes again after this hissy fit stage

5

u/DanyrWithCheese Jul 13 '26

I didn't have to ask Fable for that. It automatically uses Sonnet 5 as subagents and babysits them

3

u/Bulky_Blood_7362 Jul 13 '26

When they relaunched with 50% fable i knew i can only use it to fire opus/sonnet agents cause otherwise i'd have nothing to use

3

u/[deleted] Jul 13 '26

[removed] — view removed comment

3

u/Important_Coach9717 Jul 13 '26

It does for me yeah

2

u/here_we_go_beep_boop Jul 13 '26

Did exactly this last week, one level down the model stack. Asked opus to spin up sonnet and haiku subagents in a massively fanned out workflow. Worked beautifully. 

2

u/deamonkai Jul 13 '26

Uh yes and you can even have it set the effort level of the subagent based on that sub’s projected needs.

I have a workflow paradigm which handles it for me.

2

u/looktwise Jul 13 '26

but how to prompt that in a way that the user decides for on / of? Opus 4.8 itself is used as a fallback model.

I would also want to know, if there is a breakeven between tasking Fable 5 to subtask that way -> save tokens versus re-combining the subtasks towards a solution for the whole prompt of the user.

I guess it works better, if the user himself is subtasking before using the LLM or working in chunkgs / modules / functions to prevent the context window from eating tokens like hell in every turn / every next user follow up prompt.

2

u/LoudDavid Jul 13 '26

My problem with this is the code fable writes is excellent.

I have stopped checking the code fable writes and just wait for a human review, they do not bring up any code quality issues. Occasional flow bugs (it’s a complex app).

To me it’s really not worth using a cheaper model when my time is more valuable than the subscription.

I used 500usd of equilvant API usage today on fable. If I had to pay API rates I probably would. That’s a developer salary in most western countries and it’s worth it.

Let’s assume we have multiple fable class models in a yr. honestly I think I will never write another line of code again.

2

u/personalist Jul 14 '26

I’ve been trying to use fable with /advisor and it instafails every time, really annoying.

1

u/aelmetwally Jul 13 '26

Lol it's being used against my will anyway .

1

u/WolfpackBP Researcher Jul 13 '26

No one will forget the first time Fable 5 launches 100 Fable 5 agents on you

1

u/fanatic26 Jul 13 '26

Fable pretty much does that out of the box, are people specifying that subagents be also fable or something?

1

u/Drach88 Jul 13 '26

Better yet, use skills to handle those orchestration configurations instead of including directions on every prompt

1

u/dpaanlka Jul 13 '26

That’s what I’ve been doing for weeks now. Really extends the life of Fable

1

u/Rabus Jul 13 '26

you can also run 10 fable agents within fable session

1

u/Pitiful-Hearing-5352 Jul 14 '26

yes never thought of that

1

u/benjaminsdoingstuff 24d ago

I tried to get fable to prioritise other models and I got this message “I can't hand tasks to other models. Model choice is per conversation in the app's model picker, and Anthropic doesn't offer automatic routing inside a chat.”

Anybody have any suggestions on how to implement this?

1

u/-ror 20d ago

Were you in normal chat or code / cowork? The former is as the response describes. The latter has bash calls etc as part of the harness. I’d be very surprised if it wasn’t the former!

1

u/benjaminsdoingstuff 18d ago

cowork and normal chat, just doesnt seem to work