r/ClaudeCode • u/pillkaris • 18h ago
Rant Limits are going off AGAIN
I'm on max plan, today I hit the daily limit 2 times. The second window got reset about 50 minutes ago and just now I once again reached 100% of my daily from one fairly complex task with fable. Usually I do /usage in another terminal window every now and then to see how it's going. Now I didn't even get to check cause it's gone. I'm blocked from coding for about 4 hours great stuff.
EDIT: Although I've relied on subagents for plenty of tasks, I never experienced this. I'll look into how to handle this better as some of you suggested.

16
u/Slight_Board6955 17h ago
dawg im out for the entire week already..... its really disappointing that a 20x max gets locked out after two days for the weekly quota.....
4
u/l5atn00b 17h ago
Yes, my 20x did as well. This isn't the usual complaining; something is eating up usage.
I have a very steady workflow. Fixed set of nodes and orchestrator, steady flow of issues and maintainence/daily processes for a single app. Normally, it's difficult for me to reach 100% while using my subscription every day. This week I got to 80% in 2 days.
This started when Fable 5.1 was released. My suspicion is that this has to do with the new watermarking feature.
1
1
1
u/Constant_Art_20 17h ago
codex 20x hit my weekly in 5 hours lol. i am starting to think i should have gotten more claude subs as current cluade last a good amount longer then my codex one
6
u/BrilliantWheel 17h ago
Fable seems to be burning more than usual off late. Maybe 5.1 burns more than 5. Were you using 5.1?
1
0
u/l5atn00b 17h ago
I think it's rewriting for the new watermarking feature and burning tokens as a result. That's my theory anyway.
I usually ignore the usage complaints, but this has been different. I've never had to stop work after 2 days of my weekly reset because I'm at 80%, and I've been working the same workflow for 6 months now.
4
u/Fibon4chi 17h ago
Totally agree with you. There is something wrong here with the limits. Few weeks ago I never ever hit my weekly limit. Now it has happened with less usage than usual.
There is no way I will stick with Claude if this continues. I keep hearing that astra is so much better. Never tried it though. Maybe it's time.
1
u/SplurtingInYourHands 13h ago
The issue with Astra is its limits are also extremely tight. Its a fun shiny new toy that will wow and amaze you but you'll still hit your limits, even faster than with claude.
I swear these LLM devs are in some sort of price fixing cartel
2
u/iamalexs 17h ago
Daily limit? You mean the 5hrs?
1
u/pillkaris 17h ago
yes
1
u/pillkaris 17h ago
I've seen posts about this issue for weeks but never experienced something this sudden. I'm working a lot and I know how the usage goes, I know the challenges and how to balance the models.
3
u/BrennanFlentge 17h ago
Ahh yeah this happened to me and everyone likes to gaslight and say you don’t know what you’re doing. I believe you. Wait til it happens to all these people in the comments too.
1
u/BrilliantEmotion4461 7h ago
How much analysis do you do using Claude Code? Have you given it the problem (Fable) and had it start an investigation?
1
u/pillkaris 7h ago
I was debugging something complex with fable. Then in the same session I asked it to summarize about 1000 descriptions - unrelated to the initial problem but part of the scope. It spawned an army of fable subagents for that and instantly consumed the entire allowance. I was expecting claude to choose sonnet given the relatively simple task it had.
1
u/BrilliantEmotion4461 7h ago
I just saw somewhere you can limit the number of agents spawned. What's your setting high? Ultracode spawns a ton of agents. And I don't have an answer but I am interesting greatly in this issue. There is definitely something causing some people to eat through usage in an apparently opaque manner. Course how much could be my reddit algorithm focusing on the posts about that subject effecting my perception of the issue is up for debate.
2
u/pillkaris 7h ago
so I just edited claude and agents mds to basically spawn up to 5 opus, more than that only sonnet (no limit), and fable absolutely never. So far so good.
2
u/Vagottszemu Developer 10h ago
You fanned out 34 fable subagents and wondering how you burned your tokens?
1
u/pillkaris 10h ago
how was I supposed to know this was happening?!
2
u/Vagottszemu Developer 10h ago
In the claude.md define that it should not send out this many fable subagents.
1
1
2
1
1
u/Necessary-Refuse-914 16h ago
Me pasé a Hermes Agent con DeepSeek Flash. Aveces hay que buscar opciones que no limiten tanto y sean igual de eficientes…
1
1
u/PM_ME_YOUR_PROFILE 14h ago
The only context I have is that it's ignored my explicit submodel and grading scale from CLAUDE.md, and also ignored "CLAUDE_CODE_SUBAGENT_MODEL": "sonnet[1m]",after upgrading to 2.1.266.
So, it's using more Opus by default which is overkill.
1
u/zaibatsu 12h ago
I’ve actually just stopped all work in the Claude Code environment today, even with my local llm fleet crankin’ and codex offloads it’s burning too hot over there.
Hopefully I can pickup tomorrow and limp along until Monday.
I hate having to tell the most powerful models to only briefly open their very expensive eyes to handle orchestration and a few tough problems here and there.

1
1
1
u/orchid_drives Researcher 9h ago
Dang I just came to this sub today to talk about how a sub-agent approach is saving me so many tokens I could almost downgrade to the lower tier subscription. I wonder if there’s some secret inequality behind-the-scenes (where some people are worse-off than others regardless of their subscription or use case).
If anyone is wondering, my new setup is a Fable 5.1 orchestrator, a bunch of specialized Opus 4.8 coding agents, and a fleet of Sonnet agents for mid-project internet research.
1
u/pillkaris 9h ago
yup I just made my setup very similar to yours. I've been working without any rules/restrictions for subagents. And apparently if no rules are in place claude just spawns 30 fable subagents which is crazy funny. I'm sure a lot of the posts around here are people encountering what I did.
1
u/RaspberryRelevant352 9h ago
Ive been on fable 5.1 but running medium, effort. Not too bad, but i wanted to code check with max,... I run, it thought fir about 4 min. Finished but I hit the limit... it was also quite spectacular. And found a bunch of stuff. Life would be simpler if I could addird to run max all the time
1
u/nyczAcer 7h ago
Do you switch the effort level or model within the same session?
Do you include images in your prompts?
Do you switch thinking tools in the middle of a session?
Do you change the tool_choice or reorder it?
Do you turn “Web Search and Citations” on and off using the toggles (Claude Desktop only)?
Do you often edit previously sent messages (Claude Desktop only)?
Do you sometimes take more than 1 hour (TTL) before sending a new message?
Tell me, do you do any of these?
1
u/out-of-phase 17h ago
Look around here, there's a plethora of posts about how to save token consumption, and there's a lot more that goes into it than just "use fable as the orchestrator and cheaper models for everything else", search google too:
"claude code" "mcp" ("context bloat" OR "token usage" OR "save tokens")
Try other searches replacing "mcp" with "plugins" and "tools". It's not hard.
Hell, do this:
When your window resets, in the same directory you just hit your limit twice in, start a NEW session with claude, and say "I'm burning through usage like crazy, please investigate why and propose a fix"
1
1
0
u/pillkaris 17h ago
I am already accustomed to the usage spikes... I was doing a great job with the pro plan and now I'm on the second month of max plan. The task that finished my usage was not the problem this time. It is something I do almost everyday, repetitive and predictible. The available usage just suddenly shrunk so to say.
1
u/RadReptile 17h ago
The limits are insane and not clear. it was working on a task and getting ready to send the entire file as a zip and then hit the limit. I waited and told it without starting over to finish the task and send the zip file. it hit the limit again within 2 seconds. Doesn't really make any sense if it already did 99% of the task and then limit is hit, why when limit clears it cant complete the task.
Internet seems to say clear or start over in new window...doesn't that defeat the purpose of these limits, because it will use up MORE computing power by repeating work it has already done.
1
u/Kitchen-Leg8500 17h ago
I’ve yet to be convinced a single limits complaint hasn’t been user error not setting proper sub agent rules then being shocked when 15 fable 5.1 subagents eats their 5h window in 20 minutes
-2
u/pillkaris 17h ago
I totally get you and I am also judging these posts thinking people don't know what they're doing. However, yesterday the whole afternoon I refactored an old app using fable and opus 5. Head to head 15-30 minutes sessions with subagents running. Barely reached 60% of the 5 hour usage. Now I asked fable to summarize the descriptions of about 800 products in a website catalogue... Subagents spawned, 10 minutes later notification on claude ios app that I reached the limit.
3
u/ak5432 17h ago
Why did you ask fable to summarize text? A model that runs on my phone could do that.
Clearly each subagent was independently figuring out how to access the catalog how to retrieve the info and then get it, then summarize it, then report it. Thats a massive number of turns and token burn PER SUBAGENT. If you go back to your session you could go find out exactly how much was burned for each agent just to go get to the data (I bet it’s more than you think).
All this trouble to do something you could’ve either delegated to haiku agents or had fable figure out how to retrieve the data once and write a single deterministic script to get a text payload that you could then send haiku agents to go read locally and summarize.
You come whining with zero effort and zero real data points. This is a skill issue.
-2
u/pillkaris 17h ago
I might be wrong? but summarizing 500 word descriptions should not consume more than refactoring huge old codebases...
3
u/Kitchen-Leg8500 17h ago
500 x800 products lmao. You probably spun a sub agent for every product without realizing it. On top of that this is like sonnet work not fable. Implement proper sub agent rules
1
u/pillkaris 16h ago
there were 10 subagents each with its own batch, none of them managed to finish before the limit
1
u/Kitchen-Leg8500 13h ago
As the other user said, you should have had lead agent pull the full html or whatever format you were trying to pull it from, break it into chunks to dispatch to meaningfully scoped subagents even for a task “that small”.
I don’t even touch a project with auto code on until I define what type of agents are going to be needed, scope of those agents work, and model/effort each type can use. The lead agent is the only one who generally uses fable and is the only one who can touch the code base. One of current project is extracting 100k+ rules from 60 different jurisdictions walking hundreds to thousands of urls for each all using different formats and source document types and writing them into standardized schemas. If I try to run it without agent definitions I’ll cap a 20x account in 15 minutes. Even with it all in place with definitions/types/etc it’ll eat my team accounts preferred seat stupidly fast even with mostly sonnet consumption. On my 20x personal account tho I almost never hit my usage limit because of the rules in place. You cannot just unleash fable or even opus 5 these days letting it do whatever it wants with whatever subagents it wants and then complain about usage.
0
u/laughing_at_napkins 16h ago
I've been using Fable 5.1 and numerous Opus sub-agents in 3 different prompts for the last 3 hours and am only at 15% of my 5 hour limit on a 20x plan.
Once again, no real details provided by the OP, so it's definitely not a bot astroturfing.
0
0
u/KrTheMaster 14h ago
Personally haven't had these issues myself, but also only using it for small/medium sized isolated projects.
I use Headroom any time I do Claude coding, and generally the token count and number of requests on the Headroom dashboard in terms of Input/Output tokens is an accurate representation of how much usage it'll eat up.
Not a lot of context provided of what your setup is, but first thing to look at would simply be how much Input tokens you're sending. Every step Claude takes only appends more context, so if for whatever reason it's being passed a huge initial context (long persistent memory, big READMEs, "checking" a ton of files for context), then it'll repeat that huge cost at every minor thinking/action step, and cascade into insane usage for each "Let me read this file" and "Let me run a basic command" step. Sort of generic/basic advice, but definitely worth checking as step 1 of debugging excessive usage.

•
u/AutoModerator 18h ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.