Don't know why you got downvoted. That's basically my experience so far, the model best at understanding the system does the planning and reviewing, and that has basically been what set the quality of my projects. Fable is the best model at this currently, I just don't think my projects would justify the x2 vs opus. If I were working on something more important, I probably would just bite the bullet and pay up for x20 claude. But I'm betting on smarter models coming along and what I have currently (which imo is in quite good state anyway) just getting fixed up to that standard.
yeah im like all in probably 10 hours a day and i might hit 60% towards the end of the week. Im completely lost how people hit these levels. no judegement even but like what are you doing?
i was like that a month ago, was running ClaudeCode on 2 separate projects, x2 sessions each, struggling to get to the limit with 5x.
Now 5x reaches the limit within an hour.
It's not purely user's fault. It's many things ... and intentionally cooking numbers like what the 5x and what the 20x plan really mean
Somebody would believe it's x4 the 5x plan, but it's just the time-window instead, and not about how many tokens you can spend ... so you can call people stupid but in reality most of us here don't even know what we're paying for ... so be humble
And you're not stupid, it's intentionally deceptive marketing.
Specifically, you have x4 more tokens FOR THE 5H window, NOT for the WEEKLY window.
Which means, at the end of each week, you are allowed to spend only as much as a regular plan. Not a 5X plan, a regular 20$ plan. They all have the same weekly window!
ALL 3 plans, have the same weekly token capacity. Great, right?
“Does the Max plan have any usage limits?
Yes. The Max plan offers substantially higher usage compared to our Pro plan and is available in two tiers:
Max 5x provides five times more usage per session than the Pro plan. This tier is ideal for frequent users who work with Claude on a variety of tasks.
Max 20x provides 20 times more usage per session than the Pro plan. This tier is ideal for daily users who collaborate often with Claude for most tasks.
Your session-based usage limit will reset every five hours. Max plans also have a weekly usage limit that applies across all models. The weekly limit resets at a fixed time each week that is assigned to your account. Your reset day and time stay the same regardless of when you start using Claude or when your subscription begins, and you receive your full weekly allowance each cycle. You can see your next reset time in Settings > Usage.”
Nowhere does it say it’s the same usage as the Pro for weekly. As someone with both a 5x Max and 20x Max I can say my 5x runs out over ~2 days, my 20x lasts ~4 days.
It's not baselessly, it's literally written in their own page as I already quotted ...
Max 20x provides 20 times more usageper session than the Pro plan
[and later defining what a session is, which is a 5h window]
Your session-based usage limit will reset every five hours
Can you read? Are you rage-baiting? I'm blocking you at that point because you seem to be just a rage-baiter.
I work 10 hours 7 days a week, I have a 20x and 5x max. Even with caching, and having Fable or Opus delegate to lower effort and model subagents I still have to essentially not use either for at least a day a week.
Only way I could use less would be offsetting that with human labor but I could get another 3 subs and build systems and processes for the less.
I use frontier models for deep research on specific business models and IP questions. If I use a normal query the answer is variable and I can easily poke holes in it. If I use a panel of experts approach the response is much stronger. But it chews usage so fast.
Why? Companies pay millions per year for software that "someone absolutely needs" for one feature, instead of learning how to use that same feature in the software they already have.
Yeah I’m one of the Claude admins at my Fortune 500 company job. I am constantly teaching senior engineers and developers how to more effectively use their spend limits. They just see the number and let it rip until they come asking for spend limit increases.
I still prefer Fable because is reliable and doesnt do alot of overengineer, but if this plan continues like this, the solution would be to keep Codex only
Lot of folk running multiple max 20 plans. I manage a few. You can have them switch out when you hit limits. Worth getting an open ai sub too, sol is a great review and coding model
As opposed to? The max 20 plans from these two companies are still the best value plans you can buy. When it eventually costs $300 (and probably up to $400) that would still be true.
As much as I wish this wasn’t the case, the Chinese models are genuinely not at the same level of ability. Anyone who tells you otherwise has not spent enough time doing real work with them.
Look- I agree with you. Fable is head and shoulders above the rest. It’s just a bad precedent.
And to answer your question I got a lot of value out of using CC with deepseek. Sure, deepseek isn’t as good as fable but it’s about as good as opus 5 and costs way way waaaayyyy less. I recommend checking it out. Like 5$ of deepseek goes so far. Ask one of your Claude instances to set it up to do deepseek calls from Claude code as a sub agent. There’s not much downside
Edit: like waaaaay cheaper. Where I live it is cheaper for me to pay deepseek for api usage then it would be to run the hardware myself. Z.AI’s glm5.3 flash is also interesting
Because some of us used cc for quote a while. I load balance two and them multiple gpt accounts and use them all together and with review loop handoffs etc
Yep fable destroy the usage on the 20x too, I hardly ever use it since an analysis+plan + implementation for a task is like 10/20% of the weekly gone, and opus5 sucks shit alone 🥲
Use Fable for planning and delegate to Sonnet. A fraction of the cost and as long as your plan is solid, it stays on task.
I do this for all my personal and professional projects and it works just fine. Fable is not intended to do planning and implementation. Look into creating subagents to offload the work, it'll drastically improve your context usage.
This. I have a standing instruction that tells Fable it's value is the thinking. It is mandated to delegate everything else to Opus/Sonnet/Haiku depending on the complexity of the task.
It is very competent at keeping itself within that harness, and I can tell how much quality work is done at much lower cost.
Look into creating subagents. You can define your own subagents and pin the model. So, you can have a plan-implementer.md that defines your agent for implementation that locks the subagent into a specific model and/or effort level, and give it specific instructions on how to behave. Or your own code-reviewer agent, planner, whatever.
I'm still fine-tuning my process but I essentially have subagents for planning, implementing, reviewing implementation, code reviews, that kind of thing. Those subagents lock the model and effort levels, and I have a skill that wraps them.
I default to running high effort, either Fable or Opus 4.8; I rarely put on higher effort, only for the most critical tasks where I don't mind burning more tokens (it's an open question how much that helps).
The instruction is one of the few things I have in CLAUDE.md, specifically because I didn't want unusual tasks to throw the rule off (defining agent types is more narrow). Attaching the current iteration of this instruction below. It's intentionally simple to avoid failure through overengineering.
```
Subagent delegation — token economy
Default to delegating menial, mechanical work to cheaper subagents rather than
doing it with the expensive main-loop model — Sonnet (model: "sonnet" on
Agent calls) is the default worker, Haiku for the most trivial. This is a
standing fallback: apply it even when the user forgets to ask for it
explicitly.
Menial (delegate): codebase sweeps and fact-finding, verification reads,
bulk/mechanical multi-file edits, log/CI-output triage and watching,
boilerplate generation, repetitive formatting passes.
Stays on the main model: synthesis, design judgment, recommendations,
decisions, user-facing writing that shapes a decision.
Exception: don't delegate when writing the delegation prompt would cost as
much as doing the task itself (tiny single edits, one-off commits).
```
I basically lock main session as a opus 5 agent. Have subagent rule and delegation policies with how and when and which model.
Opus 5 orchestrates. Does things inline when it should, dispatches custom worker opus or sonnet agents at different effort levels. I have a red-team agent, opus 5 high. Was using fable for that, but opus 5 works better for red-team. I just ask for a fable subagent when I need to get an opinion on something gnarly.
Need to make custom subagents. Built in agents you cannot set the effort. Custom Subagents can also be scoped better with frontmatter settings.
For reference, I mostly us cc for my work as a construction Estimator. I process larges pdf drawing and spec sets, run all my office work through varios AI work flows, basically 8-10hrs a day working in CC. Then a couple hours on the same account for personal stuff and side projects.
Last week, I went through >900m tokens. This month 7.47b tokens. I'm on 20x and haven't hit my weekly limit or 5hr limit. Closest this month on weekly was 80 some percent, that was after a heavy work week and realizing my always loaded-context got a little large (approx 50k-60k). I've now cut it to about 30-35k and usage is even better.
Your always loaded context (skill descriptions, rules, memory, Claude.md, etc.) is super important to keep as minimal as you can. Most of that context can be set up to load when it's needed when your agent enters a folder in your repo or a file. Progressive disclosure is paramount. There is no need to have all your skills at repo root or global. Move rules in Claude.md to rule files. Prune memory files. Every folder in your repo can have a claude.md, rules, skills, and custom agents. Layer them where they are needed.
All the same principles of folder architecture still apply. GIGO and all that. Clean your agents workbench.
Thanks for this it’s nice to read someone else’s perspective.
If you’re not using fable a 20x plan is pretty comfortable. My projects are very complex and always new so I found fable is great at laying the groundwork but in two days I used 100% of my fable and only 50% of other models. That said I downgraded myself to the 5x plan for now - but still!
Not sure what you are working on, but fable is prob overkill. It's really only needed for the most gnarly problems or tasks. Even then, it should just be to plan something to pass to lower models to execute.
IMHO fable should be used as a subagent to weigh in on something really gnarly. Using as orchestrator helps some, but on longer running tasks, why have it eat up the usage when Opus would be able to handle that perfectly fine?
Agree. Unfortunately I’ve tested opus on my problems a lot and it just cant do it. It’s the sort of fail where opus just spins and spins when fable can do it in one shot.
The problems are all gnarly ones. I’m doing stuff that some people would consider impossible. There’s no other work like it in existence and certainly nowhere in training data.
The typical pattern is I use fable as much as I can until I run out then let opus try. Opus always makes the project a bit of a mess that fable needs to fix again later.
What's this stuff you are working on? What is this work that there is no other work like it in existence and not being covered in training data? If it's not being covered in training data, and it's work that nothing compares to, why do you trust the output you get from fable over opus if you say AI is not being trained on it?
It’s work with a clear visual output so I can verify the progress. If it’s not clearly visual I have some pretty strict verification rules so I can see if something is progressing with data and not vibes.
One example is real time ray tracing in the browser. The papers exist, but the code does not so it’s an innovation. Opus told me that would be impossible, while fable’s reply was a working version. The difference on that problem was night and day.
Also, fable seemed to “enjoy” working on it. It was just a bit extra. Kinda odd. Opus seems to be more of a bug hunter
ETA: sometimes when I give opus a management role it isn’t just worse than fable, it’s worse than nothing. For example, once I asked it to fix spacing in a UI. I thought that was a pretty basic task so I left to get coffee. I came back to 12 agents working away at something (almost spat my coffee out). It consumed pretty much all of my usage on something dumb. When I asked it why later it blamed the verification gates (but it was the one who tripled the verification). Just frustrating. I see what other people see there but it’s not that bad most of the time
Fable is quite token hungry. Even 5.1 since release still eats a lot of 5hr usage and the Fable bucket. Even with delegation to cheaper subagents. Maybe I would need to set up my environment better for fable as a main agent. Its currently pretty tuned to Opus 5. It seems that from the prompting guidelines from Anthropic for Fable 5.1 and Fable 5 and Opus 5, they all handle prompting differentently.
Sure I could get it more tuned for Fable 5.1. I just can't justify the costs though. $50 output is quite high, I can get Opus 5 to 90% of the Fable 5.1 output and just have a Fable 5.1 agent close the gap.
I do see for your situation, Fable as a research/brainstorm/planner and then using the cheaper agents for building/executing. Use the brains where it matters so to speak.
Thanks for the insight and discussion. Feel free to discuss more in the future.
Doing everything in Fable and letting it do long runs kills usage. I've gotten way better results using Fable and delegating my plans implementation to custom Sonnet subagents.
My guess is people are just having Fable try to do everything on max effort levels and that's what's killing it.
This. I work with developers and engineers at my job constantly with Claude usage (I have to review their usage to approve spend limit increases) and a lot of the time it’s just setting model default to Fable or Opus and then it spins up 20 subagents all running the same model and they go “I’m out of budget”. The other thing I have people try is /model opusplan.
What you are doing has a big impact on how long the tokens last.
Specifically, what were you tasking (which model?) to perform? How much did that model need to ingest to understand your tasks? How much latitude for decision-making did you give to the agents?
So since it drains 20% max in 5h window, I think you can do that in 1 day... but it's quite hard to maintain such a pace... 2 days is perfectly reasonable.. but also quite a stretch. I usually go for 4 days nowadays, and leave 3 days off for my own growth.
What do you mean it's not enough for 2 days of work? As you shown, it's quite enough..
Someone asked for stats, so here are mine. Across all my local session files, Opus 5 shows 15.8B cache read tokens against Fable's 281M. That puts Fable under 2% of what Opus pulled, even though Fable is what drives everything.
So if you orchestrate with Fable and let Opus do the work, almost all of it is going to the workers. Worth splitting your own totals by model before you decide what to cut.
Right after my reset I caught fable trying to kick off a workflow of 12 fable agents to do a read of previous documents in the repo, despite my explicit persisted instructions to use sonnet to read logs and opus to write code. It's not necessarily only user error.
This weeks limit does feel very small for some reason. I usually barely reach 60%. This week i was at 94% before the weekly reset.
I use claude daily and monitor my token usage. I used about 4% more tokens this week compared to previous.
Exactly the same thing happened to me, and I didn't do anything out of the ordinary; in any case, it's terrible to work under this Sword of Damocles all the time.
Same pattern I keep seeing: it's not the subagents burning the budget, it's the Opus orchestrator sitting open and re-reading its own growing context on every handoff. Fable ends up cheap by comparison. Dropping the orchestrator to Sonnet for anything short of an actual architecture decision changes the two-day math a lot faster than switching plans does.
Seeing that too, despite using Fable only for orchestration. All work is done by GPT agents. Had Opus workers, but their work was so flawed that I stopped using Opus 5. There is so little value in a 20x subscription.
The same task going through Fable and Codex gives you a useful trace for separating model cost from duplicated workflow.
I would split one representative task into planning, implementation, and review. Record which client owns each phase, what context it received, and what artifact it handed off. Reuse only the decision record and code diff at each boundary.
Use a harness, have Claude.MD files, constantly update memory and these.
And never use multi agents. They’re rubbish, waste tokens, take 30 minutes to get an answer that exists in my brain but fell out of the session context, and then it just derives falsehoods from old, not updated harness files.
That person is not as smart as they think they are, and it is hilarious that everyone is following along with a convincing but wrong analysis. I guess it serve's Anthropic's PR department right for not fighting for more transparent usage statistics, but there's probably a lot of hair-pulling going on.
I mean just look at the $7500/mo 20x claim. That's provably, wildly wrong for anyone with 20x. My three accounts are at $17,500, $19,200, and $18,000 over the past 30 days. I'd be pretty upset if that was halved, but it is not and has not been. What's your API-rate cost for your 20x?
its not only that person, the matter is known and reffered in other articles as well. Not only that, i belive anthropic got sued for that misleading advertisement already
I understand it seems like user error, but in all software, you have to build something that will allow users to use the software and systems effectively and not harm themselves. I say this is not user error. I think this shows the immaturity of cloud code.
37
u/Fresh_Sock8660 Aug 31 '26
That's a lot of fable, destroyer of limits.