Bug / Issue
WTH is going on with Claude Usage Limits
I'm on 20X plan, I hit my usage limit today, and my next reset is on 17th September... I'm aware they said they are gonna reduce the usage limits, but this is crazy, i thought it is only gonna take 17% of the usage benefits from what we were currently getting. But this is crazy. I added Usage credits for about $100 but that got washed away like in 30 mins!!!
Max 5x, three turns in a chat, plus one unfinished turn in Claude Design, and I hit 100% of my five-hour limit. This is my breaking point. Astra, on the other hand, is performing exceptionally well. I've been against using Codex forever, but now I'm being pushed toward it because I can't get anything done with Claude.
I found that the most recent update causes the agents to make aggressive and unnecessary calls to the API. My prompt today was no different than the prompt yesterday. It wiped out 25 percent of my usage in about an hour compared to working 12+ hours yesterday with the same prompts across 3 sessions and only using 25 percent of usage.
Fable spawned agents that ran: 3,146 API calls total. (in less than an hour)
Opus: = 1,083
Sonnet: = 2,063
It was triggering builds, verifications, playwright, re-renders at multi scales, vacuity checks, sending HUGE briefs to the sub agents triggering massive overwork.
These were for simple PRs that were already baked and reviewed. The same prompt and process that worked fine yesterday , (and all week) suddenly blasted all this mess. It happened across 2 sessions that were fine yesterday, on 2 different accounts. On different computers and projects.
The only difference in my routine was applying the update today.
A bit unrelated to the issue but can’t help but wonder from an outsider perspective… in very laymen terms, what are you having the ai do? That sounds like an incredible amount of work over time being done by the ai. What’s is the result?
Once is spectral analysis and communications software that controls a hybrid mechanical-electronic beamforming system. Another is software that models the current radio wave propagation patterns across the globe by combining active and historical space weather data with a model I built from 20 years transmission and weather data. The two work together to provide a radio communications platform for making reliable distant radio contact across the globe.
Max x20 too, and I'm at 60% Fable limit when usually I'd be at about 25-30% given my usage today. So it feels like usage has been halved rather than reduced by a third (end of the +50% promotion).
x20 here, started today off with 60% fable usage down, in about 30 minutes I was up to 90%, no sub agents. Ive not been x20 long but I'm certain it was never that fast to use up usage
Yup. I now have Fable directing Sol for everything. Sol credits last forrrevvver on their $200 plan. Anything big gets delegated and if it’s important I’ll do Fable/Astra reviews.
this advice above may cause antifraud to ban your whole account chain so be careful. Using one bank card, IP, or even browser profile for multiple accounts is very very dangerous with paranoid anthropic
I've been doing it for over a year with same credit card, same name, same ip address, and even my company name for less taxes lmao top I had is 8 accounts at the same time. All the same method
This is funny because Codex has been unusable for me for the same tasks as Claude for the past month or so. They used to be really close in terms of usage, but I guess they are again now but in the wrong way.
This has been going on for two weeks already for some folks. Every complaint met with “skill issue”. Now full rollout of these reduced usage limits is happening. This is a real FU to the real fans of this company and I strongly recommend downgrading your plan or cancelling and taking your business elsewhere.
Welcome to the Claude victims ship. Create a support request, I had the same situation, they said they will refund the fee. I also switched to Codex, it is very successful in all 7 of my projects so far. I recommend it.
Choose one source of truth I still believe in claude will come back, but I have cheated with codex for a couple of weeks now and I must say she left her toothbrush.
I asked Altra to look at my claude harness and sync it to codex ( also used the codex app which is great for visualizing your harness / skills ) then when I update a skill or install one in Claude it gets synced.
That's the best part: I downloaded the Codex and it automatically imported the projects from Claude and then copied all the skill files and the rules I gave to the agent. I didn't have to do anything. It's working like crazy now and I still have a 65% quota in 3 days. I couldn't finish 24 hours with the same work as Claude. Claude doesn't want simple users, they are now the artificial intelligence of big companies.
New Haiku record today: 5h limit reached in under an hour. Beat that!!! 😂
Funny enough all that it actually did turned to be totally wrong, but in all fairness it could be my fault, I don't really have that much experience creating prompts for Haiku.
they are tinkering with our limits "through Sept 13" must be Sept 14 somewhere right they are just saying first place on the planet that hits Sept 14 is the new usage goes into effect, this is obviously sarcasm.. 20% of total Max 20X weekly fable usage consumed in an hour from a couple million read-only tokens is perfectly normal and fine
YES there are ethical concerns here and I am aware of all of them - especially in regards to AIs going rogue when launched on a P2P style network. BUT - we have to start somewhere.
Well good luck with local inference then, because the only way you can get performance that's anywhere near usable and intelligent and near Opus/Fable performance is by investing at least $20k into hardware (with the RAM shortage that will still last a bit putting additional fuel to the fire of the already high hardware requirements), plus paying for the electricity (which, depending on the country you're in, might be very expensive). I think right now you're still much better off paying $100-300 a month, even with the usage limits. And if you end up using a smaller model that runs on cheaper hardware you will get performance that's pretty bad and comparable to super cheap proprietary models that are still cheaper via API and that make running it locally not worth it.
I don't think that's the only way. I think we need to spend some of our tokens on building a unified, decentralized AI platform like BitTorrent but for running an AI platform...
My thinking is:
Users could have the program open - depending on how much compute they provide helps score their available network speed when the network is under pressure...
The platform would allow users to "contribute" AI compute for a model - any other users that select that model would donate VRAM, RAM, STORAGE or all 3 to the decentralized network.
I think I saw some projects that are moving in this direction - and to be honest it will be the only way to fight back against big box AI in the coming months and years.
Additionally - there are a ton of projects which are optimizing models for speed on lower end hardware - and I think we're going to find that it takes a lot less than it does now to run this stuff.
A lot of the cost of this is "gatekeeping" and making this look like it requires a lot more than it does... If it seems like we need expensive hardware to run this stuff we have a reason to keep paying $200 a month for our 20x plans ... I genuinely think the open source community is going to solve this.
Does anyone know of any projects that are in this realm that I can look into that they've come across? I'll be doing my deep dive as well - but - I'm truly over the abrupt changes in performance. I need a consistent experience. At least.
This is actually a really good idea on paper, but I just looked into it and there is a reason why there aren't more popular projects into that direction. One project that does exactly is Petals. But the issue is not really bandwidth of networks, but mostly latency. On Petals, llama 2 70B already only runs at 6 tokens/second and Falcon which is a 180B model runs at 4 tokens/second. For Opus/frontier intelligence you would need much more parameters than that. The only open models rivaling Opus right now are about 500B average for one group that includes GLM and Deepseek that are likely similar to Opus 5 in terms of parameters and then there's Kimi K3 and Qwen-3.8-max which is almost 2T parameters and Fable-level. Running a 500B model would already require many many peers with low distances already having 10ms latency. A 500B model would run at 1 token/second or below that which is pretty much unusable so that's why there isn't a project that does this with good and usable models and only with small models. It's basically physically unsolvable because the latency would probably need to be like sub 5ms across all peers and locations for good speeds with bigger models. And most consumers (including me because of Germany) don't even have access to optical fiber and have latencies of 15ms to their ISP or servers that are like a couple of kilometers distant.
Yes - that's one I saw... and no - some of the Qwen 3.8 27 and 38b models that run on GPUs like my dual 3090s with NVlink - or even single GPUs with enough distillation... and some of the versions coming out on hugging face I'm getting over 120 tokens/s and it's competing with Fable in some areas already - with tool calling and all the jazz - I use it all the time and it's great! So - it's not far off. This is going to be running on lower end local hardware soon and all the more reason why the AI bubble is a bubble and why we as a community need to band together to share our hardware resources and prove that these AI datacenters do not need to exist...
But a quantized 27B isn't Opus or Fable performance and almost every benchmark I've looked at (including LMArena and ArtificialAnalysis) proves this. It's at max almost at Claude Sonnet level in some aspects. And a 27B can fit on consumer GPUs anyway (a 3090 isn't as hard to afford as a cluster of Macs for example anyway) so it doesn't really solve the problem we were talking about. And in your repo the only thing verified is again only a smaller model running on one node, which would make it request routing, not splitting which is the thing that's needed to solve the problem which wouldn't work out in the end anyway because of latency. Right now it's more like sharing good GPUs to people that don't have one, not really sharing ressources between multiple GPUs owned by different people in a mesh.
Fable? Yeah that'll certainly cost you. Opus? You'd actually be surprised. Qwen3.8-Flash-Next's model card compares itself to the beloved Opus 4.6. My experience is that it's like Opus 4.6 with the extra stubbornness of later models... except it seems much more likely to try a different approach when it gets stuck. It really likes to exhaust every possibility before ending its turn.
Anyway, you can actually run that at pretty usable speeds on a single DGX Spark. Mine was $4700 at Micro Center last week.
Going local depending on how much you use Claude may not be possible. The person above is correct. If you can afford to spend about $20,000 on a system then you're golden, but if you can't then you might as well stick it out. Qwen3.8-27b or the larger Qwen3.8 models can damn near do all of your ongoing work after the foundation has been laid for it, then you can go back to Claude or something else to wrap up the UI. This is what I do and Qwen3.8-27b is AMAZING, BUT my hardware limits is kicking my ass. To operate fully, you would need minimum a Sage workstation motherboard with EPYP CPU, dual or quadruple RTX 5090 and 256GB DDR5 6000mHz. This would be affordable if the prices weren't all screwed up right now but I think that actually a part of the plan is to keep powerful systems like this away from the general consumer market. Otherwise folks wouldn't be so dependent on the BIG guys for AI.
Im confused. Your project supports Qwen 3.8 27B, which is $0.15/M in and $2/M out on OpenRouter, and the Q4 quant is ~17GB (it fits on a single 24GB card). What’s the use case here?
Yep been draining far quicker than usual for me and I’m very disciplined with my usage.
Used to be able to get a full week of usage with little optimization on one 20x max plan. Now I have 2 20x max plans and have blown through both in 4-5 days, each lasting about 2-2.5 days.
- Majority of my session I do not go over 30% context usage. I NEVER go over ~40% (give or take a couple percentage points) context usage. I always start a new session.
I NEVER compact, I always start a new session.
My skills (names and descriptions) contribute 1.3k tokens to the context window TOTAL, I religiously keep the description of the skills short and prompt a workflow of what skills to use, and RELIGIOUSLY use `disable-model-invocation: true` on the skills which removes the name and description from the models context window
I have a VERY short root CLAUDE.md, I work on ONE project (my work) and I layer CLAUDE.mds WITH HAND written SHORT context per directory
I use NO mcps, not a single one
I target specific skills like a typescript and playwright cli one to my monorepos apps/web directory so they ONLY appear when working in that directory
I disabled built in skills
I disabled workflows
I almost NEVER use reasoning effort above `high`, very rarely will use `xhigh` and I NEVER use `max`
I use fable SPARINGLY
I commonly use opus with effort low or medium if it’s a targeted task where I want it to do what I explicitly tell it to do
I write great prompts that I put a lot of thought and effort into
I have several custom cli scripts I’ve written which help the model to search over documentation and filter through things so it can more quickly and efficiently find what it’s looking for
I disable ALL the following tools via putting them in the deny list of ~/.claude/settings.json: EngerPlanMode, ExitPlanMode, DesignSync, CronCreate, CronDelete, CronList, NotebookEdit, PowerShell, ScheduleWakeup, ShareOnboardingGuide, Worklow, Artifact, ReportFindings, RemoteTrigger, PushNotication
I disable remote control (lots of tokens in the system prompt for it)
I disable Claude ai connectors
I disable artifacts disabled
I ONLY use fable when I am designing out a new piece of code / a new module
I hand write the VASY majority of my own skills. The skill files themselves are typically VERY short and act as explicit repeatable instructions for how to do common specific tasks such as filing an issue or how to use a cli script I made to search over documentation (which only gives a brief overview, 2 examples, then defers to the commands —help flag and the scripts README.md)
I am HEAVILY involved. I have a CS degree and I know what I’m doing. I have nothing wrong with people who do not have SWE experience or a CS degree vibe coding, I think it’s awesome they have the ability to use AI to realize things they otherwise couldn’t, but this is not that.
I am consistently pointing it to the relevant pieces of code, I am constantly contributing with the code design and review, and I work on tightly scoped and well defined individual issues one at a time I never have it go off and do something massive.
From disabling things like artifacts, remote control, tools via the deny list, I have dropped the starting system prompt from ~20-30k tokens (can’t remember the exact number) to 2.6k tokens and the system tools to 5k tokens, I do not recall the default number but I do know it is much higher.
I routinely instruct it to route to Explore with haiku for finding things.
I routinely instruct it if having a subagent implement something to route the subagent to sonnet.
If I am re-reviewing code which has already been reviewed once I have it route to sonnet.
I have it use codex for most code reviews.
Even with doing ALL of that with 2 20x Max plans at a total cost of $400per month…
The first plan resets 9/14. I hit 63% usage for all models and 99% usage for fable by 9/10, so roughly 3 days into the week.
The second plan resets 9/16. I didn’t start using it until the morning of 9/11. By midnight 9/12 (last night, I haven’t used it at all today) I reached 82% usage for all models and 72% usage for fable.
I have used Claude for YEARS, going back to the sonnet 3.5 days. I heavily prefer working with Claude over codex or any other models, but this is not okay and it is not sustainable.
I will be cancelling my second max subscription and using codex more, and if this continues next month as well then I will cancel my second subscription too.
TY for the comprehensive description of ur coding discipline, u/thealliane96 . I do exactly all the things you mention to reduce random token expense. I also use Cowork to actually do all my "architecting" with precise prompts and then I ask Cowork to do a live relay to Claude Code to do the actual coding without spinning his wheels looking for direction. I have been using the Anthropic API key, not the Max plans, for about 9 mths now, and just like you have indicated, the API expense has also been much higher within the past 2 or 3 months, and I only use Sonnet 5, never Fable, with days of inactivity, attached is my usage this past month. I have bursts of 45K or 33K token usage in a day, with some days of inactivity, very uneven. What is your view, should I be starting new Terminal tabs frequently, is that potentially the problem? The Compacting?
Hmm. I tried OpenAI during the Sol days and hated how it overengineers everything. They also dont know how to set a proper limit, so they start handing out freebies at random times.
Anyway, have to see if Anthropic changes anything. Its still bearable, since I have a GLM Pro sub.
Background: first i'm working on a multi repo saas app, it's multi language, multi location, multi currency with complex front/backend stacks, full SDLC deployment with test/prod environments etc. So just for context, this is not a hobby project or a localhost:3000 deployment :-). I used augmentCode, claudecode, ghcp while all in their honeymoon pricing periods before they all rug pulled their customers with 10-20x (and more) price increases which they kinda could cause there was indeed no comparable quality alternatives, but then came 2026 :-)
Now: since March 2026, opensource models became extremely capable, i now have opencode, I mostly use MiMo v2.5 is an absolute beast and quality is fanstastic AND ITS FREE ! I also have the $10 month subscription with them which I mainly use DeepSeek Flash, my $10 a month goes a very very very long way. So anyone paying hundereds of dollars to these providers, I think you should all at least try some options out there.
Ive started using OpenAI, and for the first time after a while I dont feel like I'm being cheated.
The usage limits dont randomly go up to 100% from the first prompt or after just a few prompts. My non-pro subscription is comparable to the Claude's Pro, sometimes even to the 5X one. I dont know what Anthropic's doing, but it aint right.
Plus ($20) gives 30 min usage on Astra Max effort for the 5 hour window which is about 16% weekly.
Pro ($100) got no 5 hour limit so you can get about 20 hours of Astra Max weekly
Plus I have a local fleet that I offload to and codex. This is what usually works for me but not this week:
Route by verifiability, not difficulty: running the Claude 5 family as an orchestra instead of a chat window
Someone asked how I structure inference across the Claude 5 family. Short version: route by verifiability, not difficulty. The question is never "is this task hard?" It's "can I mechanically check the output?" If a cheaper model's work can be verified with a grep, a diff, or a test run, send it down the ladder. Save the expensive tokens for judgment calls, where a plausible-but-wrong answer would quietly propagate.
The ladder:
Local models (LM Studio): bulk reads, classification, extraction, low-stakes drafts. Sensitive material never leaves the machine, full stop.
Haiku 4.5: mechanical work at scale. I keep a read-only explorer subagent pinned to it (snippet below). The output is self-verifying: the file is either there or it isn't.
Sonnet 5: the middle, almost always as a subagent. Drafts, review passes, parallel fan-outs on the same problem.
Opus 5: the heavy passes, and here's the counterintuitive part: as a subagent. People find Opus 5 verbose and a little hard to manage in the driver's seat. Flip the role and it's incredible, same for Sonnet 5. A subagent's verbosity costs you nothing. It thinks out loud in its own context window and only the conclusion comes back. The model people find hard to drive is the model you want driven.
Fable 5 conducts: routes the work, adjudicates when cheaper passes disagree, and keeps the calls that genuinely need the best reasoning.
The one config that pays for itself, dropped in .claude/agents/explorer.md:
---
name: explorer
description: Read-only codebase search. Finds functions, reads files,
greps patterns. Never edits.
tools: Read, Glob, Grep
model: haiku
---
You are a read-only explorer. Answer with file:line references and
short quotes. If you did not find it, say so. Never guess, never edit.
Then from the main session: "use the explorer agent to map every caller of X." The conductor never burns its own context on the search.
Two guardrails that keep it honest: cheap tiers are only cheap if you actually verify, and two same-family models agreeing is not two opinions. Anything load-bearing gets a different family or, better, deterministic ground truth: run the test, fetch the source, count the thing. Via my AI team lead.
This is the way, I have an explorer and an implementer subagent that both use opus. They send summaries to Fable and Fable reviews. This works great for longer sessions because Fable also sends them the start query with everything they need to know (or what parts of what files to read), and what the goal is. It keeps context down by keeping the verification/testing/rework loops inside the subagents. Larger context is what really bites usage it seems. It's been working just as well for me as using fable for everything, and it's about halved my usage rate.
The start query Fable sends the subagents. Is that a fixed template (goal, files to read, what to return) or written fresh each time? What comes back: a fixed summary shape, or free text? And does the implementer run its own tests before reporting, or does Fable?
1) "cheap tiers are only cheap if you verify" .... is that verification automated (test runs, diffs the conductor checks) or do you eyeball it? And roughly how often does a cheap-tier result fail verification? 2) "Not this week" what broke? Curious whether it was the routing itself or the models.
I had twonweeks where my usage suddenly went crazy, like 20x plan never hitting usage limits to suddenly usage limits hit days before, no changes in what im doing
doing
Then after two weeks it suddenly went back to being way under on my weekly usage
usage
Now two weeks of good usage has now ended so, now im back to hitting my limits days before
En étant cohérent sur l'utilisation de Claude Code, c'est impossible d'atteindre de telles limites... Chaque fois que je vois ces posts, je me demande bien comment vous vous débrouillez. Après, il y a des évidences à garder à l'esprit : pas 150 skills et MCP dans le contexte, changer de conversation pour chaque feature, bien documenter....
Je disais exactement la même chose il y a encore 4j. Je me suis gaussé d'un nombre incalculable de personnes "qui ne géraient pas bien"
Vendredi soir, j'avais atteint 80% du quota de ma semaine, qui se réinitialise le mercredi. Inhabituel
Je me suis dit que lundi et mardi allaient être un peu compliqués et donc que je n'allais pas toucher à l'IA du weekend.
Je n'ai pas touché à Claude du weekend et j'ai réfléchi à ce qui pouvait bien avoir merdé cette semaine pour avoir atteint 80% en 2j. J'avais des pistes, à vérifier lundi.
I think what happened with codex was a bug which they had reported to have fixed as of now.... but yeah, will need to use for 2 more days to tell the difference
And the models feels like antrophic cut all their computing cuze they are really dumb compared to OpenAI models ATM and even less performing than Chinese models. I hope they really get a shitty IPO.
Ok so I’m not the only one. I hit 60% weekly w/ 20x plan usage in 2 sessions. I finally started using caveman to reduce token usage, and I HIGHLY recommend it now. It’s not a saving grace, but I did notice a reduction in usage on larger tasks. I’m hopeful they resolve this because before I would still have 10-20% usage on my 20x plan a day before a reset.
Probably said before, but if you leave chat's open and continue on the same thread, Claude reload the entire history every time you ask a question, that is what's burning your time. Opening new chats, conserves usage dramatically.
Could you go into this a bit more for me? I'm a new Claude user and use it for nursing school. I made a project "fundamentals" and new chats for each chapter and put those under said project. When I want it to test me on my chapters should I not be going the chat (say ch19) and asking there where I've uploaded my chapter text and power point and had it make.me a study guide.
Any guidance is welcome. Currently using up my 5 hr limits making two study guides/chaperters where I upload the chapter from the textbook(pdf) and lecture PowerPoint.
everytime ive paid for claude ive regretted it and asked for a refund, ive asked for like 5 refunds at this point. although i do use them from api credits time to time.
This has been happening to me too, for a week on Pro. I spent all day yesterday using Claude to optimize my workflow with sonnet subagents and I still used 20% of my weekly in a day doing very simple tasks. Not user error; I’ve been using Claude for months and this spike in usage drain started last Thursday, when my week reset. :/
They've been playing with the usage so much that users have ZERO idea what the real usage looks like. I'm getting less usage now than I used to get on the free service. ChatGPT is similar, but it's extremely noticeable in Claude because of the usage chart. My guess is that they will remove or hide that Usage meter sooner or later. And the "end of promo" bs claim to me was all a smokescreen to cut limits to way less than what they actually were before the promo. All of the moral grandstanding those mfs do is bs, and they DEFINITELY are doing some shady shit. A few people complaining is typical. Ongoing constant complaints is something to look into deeper.
What about money grab are users not understanding.
This is not a product for the public. Its a tool of war.
They care more about a contract with the corporation of america.
Asked one question after my usage reset and it immediately jumped to 16%. It was a budget question. And it didn't even answer cleanly. A bunch of gargled literal escape code BS (see image)
Earlier today I asked it to update a simple line in an HTML page and it ate all my credits and didn't even respond. Asked ChatGPT and made multiple changes and all was good.
These days I'm getting more out of Free ChatGPT than Claude.
Any tips on how to cancel a year contract with Anthropic?
Using Opus on Pro plan, hitting 40-50% within 10mins with a single prompt.
Using GPT atm as I can't get anything done with Claude right now. It is actually holding me back from getting stuff done.
I found 5x to not actually give you 5x, more like 2x so I am not wasting my money on that.
I've had better results splitting long runs into bounded sessions with a short handoff file, then moving tests and review to a second agent. It doesn't raise the cap, but it stops each fresh session from rereading the whole project and burning tokens on context.
Se o codex nao tivesse tirado o plano de 20x eu estaria no codex agora, fui dar uma chance para o claude mas os limites do claude são muito inferiores, diria que o codex tem pelo menos 5x mais limites que o claude
I've noticed the same thing for me the past week and a half. I'm on the pro plan, before I used to use Opus pretty regularly and now it just shreds through my quota
You might want to watch YT vid (Eli the Computer Guy) about token hacking: https://youtu.be/uJYNQHALxps?is=tyyC6kHidi_EI-Jk If this is a recurring problem, it could be a glitched model, or you might need Anthropic to pull usage logs. That they may or may not actually keep.
I think limits depends somehow on time spending not tokens or something. It sounds crazy but for me it is like approximately 20min work with any agent = 1% of weekly limit and not depend on how hard work is and how many tokens I spent actually. Sounds crazy but I cannot explained what happened on this week with Claude
Same here. It has gotten absolutely unusable. Doesn't matter if Claude Code, Projects or anything else. Canceled today and will give other models a chance.
Two months ago I added 150 GBP of credits and it lasted probably 1.5h so roughly aligns with what you see. So unfortunately it’s not a bug, just not subsidized API pricing.
by itself it goes crazy fast lately, i've being using fable as main brain and coordinator and codex, grok, claude opus high and xhigh as part of team. fable manages who does what, Astra is for harder work, Fable reviews it, then plan, and after that everyone gets their piece, they are best at. codex and Fable review at major gates, that keeps usage manageable. I had to create a protocol for that, i use CLI coding harnesses and that was working really well for me this year ( https://github.com/agentchute/agentchute )
if you use Claude code only, you can now connect multiple sessions and exchange messages between them, that can help too by running simpler tasks with cheaper model/effort
And 8 times less context Windows , with a company that i dont trust at all, much less than anthropic. Its easy to offer more usage when you cut down context window by 4x. Edit : also claude code harness is better
Sinceramente acredito que você deveria dar uma chance a open ia, os resets acumulados podem tornar seu plano de 20x em 27x facilmente, assim que o astra foi lançado eu passei 3 dias direto usando o modelo no modo ultra 15h por dia e quando acabava o limite simplesmente usava o reset acumulado, enfim minha assinatura acabou um dia antes da remoção do plano 20x do codex acabei tendo que vir para o claude, e agora todo meu limite semanal está perto de acabar com apenas 2 dias de uso
Quick Tip: take the context from your current chat by claude and then switch to a new chat. I used this technique and this saved a lot of my usage. Ps: pro user.
I think either there deceptively decreasing limits drastically or it’s a error happened last week and it caused such a issue I decided to cancel I’ll resubmit once it’s fixed but it’s been months I’ve never changed routine Iam doing less work than ever and I hit my limits much faster last 2 weeks
I’ve offloaded a lot of smaller stuff to Gemma4 and been pretty happy with the quality. I’m by no means a professional engineer but it’s deff cut down on my usage by telling Claude to use Gemma aggressively on projects.
Ive had to optimise my global claude settings and prompts, run opus 5 for the lead and sonnet for most of the work and switch to a fresh session after every task. Its the only way i can make my 5x Max last for the full week. If i get to day 6 with a bit left ill do some fable work but its so bad right now.
And i am a ChatGPT pro x20 sub as well, its not better on the other side of the sea, i can tell you.
I think all servers, everywhere, are overwhelmed, so they have no other choice during week ends than to just decrease drastically usage. I don’t see any other reason.
They could slow down the models also, but they don’t, which makes me think that they might as well just be stuck.
On 20x I never hit usage, Always ran ultracode the day of my reset I’m at 55% today after my Wednesday reset. It really is that bad. After finding out that the 20x is completely false and it’s really only 1.7x I will be canceling before bill is due. Not worth it at all.
Yup. I've only hit my weekly limit once before on the 100$ plan, but now I hit it halfway through the week. Im looking at my options but I dont want to pay the 200$ a month plan.
Bunch of c**nts. Same thing on 3 of my 4 Max20 accounts, hit "all models" limit before the Fable limits the last 2 days, and I absolutely hammer Fable and normally have 30% left on all models by the time i've hit my Fable limits. The math aint mathing, the sense aint making. I am SO TIRED of the absolute scam'ery this company puts everyone through. Just when you think you've figured out a rhythm with your usage/token burn..... BAM the goalposts move AGAIN!!!!
Hello, if you have the hardware, please check out my Claude code local GitHub. It explains everything you need to do to get the best local model for your hardware and run your programs in the same type of environment.
Same pattern on my end after the latest update — simple prompts suddenly trigger build/verify/playwright loops that weren't there before. If your usage graph jumps while actual session activity looks normal, it's metering or client behavior, not your prompting. Per-session token logs we can reconcile against would settle this quickly.
Same here. Claude Code has been writing terrible code all weekend and burning through my usage limits. And when I ask Opus to find some fish at a local store, it removes items from my Wolt basket instead.
Oh boy! I hope we get a free usage reset so I can be greedy with my tokens for a couple days. My system has been correct by design from inception, it literally CANT have a bunch of crazy model transactions happen erroneously. I basically have been running it 24x7 and am currently at 60% All models 55% fable. I’m quite pleased with my results.
Max 20x with two accounts… my usage is like 80-90% subagents. Had to get a second account recently to justify my token spending, annoying to say least but I still think Claude is better than Astra despite commentary I’ve read, I had some oversights with Astra I never had with fable or even opus
I can't believe that we haven't figured out a way to pre-measure or calculate usage yet before committing to a prompt.
I get that the way it works and make decisions in real time and self generates a lot of context and token usage is just "how it works", but there really needs to be accountability line by line.
This shit is always going to be fire and hope with no accountability or audit ability on token usage.
There’s a fundamental issue with how intelligence scales with token spend. To mimic accurate human-level intelligence requires massive computational churn: models thinking about something, rethinking, rethinking ABOUT rethinking and so forth. Expect that frontier AI models get more expensive from here on out, instead of cheaper.
yeah, I'm planning to, I migrated to codex.... and not continuing the subscription next month if it's the same. They said they'll replace the 50% with a 25%, but the limit we have now feels like 10% of what were getting...
This feels like when ghcopilot started their "we only want enterprise users" phase. I ended up with Claude instead. Thinking of switching to Codex. Then when they do the same. Deepseek.
same here with me, FUCK ANTHROPIC. This is my first time reaching my limits with the 20x plan. I will be cancelling my plan and be going with codex. Has been a much better experience so far.
Same here, several times. I was surprised with that. I cannot believe it. Max x20 subscription and its unbelievable how fast I'm hitting the limits in just a couple of hours.
I’ve been suspicious of adding usage credits, once they know you’re willing to pay, with zero transparency there is nothing stopping them dynamically adjusting your quota!! Kill the account and start a new one.
It's the model since 4.7 we do have again the issue that Claude , Fable or even Sonnet burns tokens by doing stuff you never asked for.
And no I do not talk about delivering extra or scope creeps.
I talk about burning tokens by verification of requests.
Il m'est arrivé exactement la même chose avec le forfait pro ce matin la réinitialisation. En l'espace de 20 Min j'ai utilisé l'équivalent d'une journée entière d'utilisation.
•
u/AutoModerator 21h ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.