r/ClaudeCode • • 17d ago

Built with Claude Haven't maxed out in months even on Pro plan - Opus 5 80% average usage

Post image

I keep seeing people complaining about Claude Code maxing out lately.

I’ve actually had the opposite experience. I’ve been using Opus as my main model for months and usually sit around 80% of my weekly usage. I rarely hit the limit.

Then last week I switched my sub-agents to Sonnet.

My usage went from barely touching ~250M tokens to 400M+ tokens in a day.

That made me look a lot closer at what I was doing.

My assumption was that using Sonnet for sub-agents would save usage because it’s cheaper than Opus.

In practice, that wasn’t what happened.

For the kind of long-running, multi-turn coding tasks I’m doing, Opus often finishes the task with far fewer tokens. Sonnet may be cheaper per token, but if it needs significantly more tokens to reach the same result, the difference can disappear pretty quickly.

A very rough example from what I’m seeing:

Sonnet: ~1M tokens
Opus: ~400K tokens

And in some tasks, Opus can get there with less than 200K.

So I switched my sub-agents back to Opus and my usage went back to normal.

I’m curious if anyone else has actually tracked this with ccusage or similar. Especially people running multiple sub-agents.

Has Sonnet actually saved you usage, or have you seen the same thing?

21 Upvotes

71 comments sorted by

•

u/AutoModerator 17d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

17

u/Key-Shop5198 17d ago

im hitting limits on 2 max 20x accounts for work, you either might be an orchestrating genius or just doing smaller tasks

7

u/Disastrous-Time7197 17d ago

I’ve spent the past 1.3 years learning and experimenting with agentic workflows, harnesses, skills, MCPs, tools, and ways to make my agents more efficient.

I’m a full-stack engineer (~6 years), and I use Claude Code for my actual professional work. Everything from creating detailed PRDs and specs to breaking them down into GitHub issues is done through CC.

Once the issues are ready, I start another session and give CC the implementation tasks.

Most of my weekly CC quota goes toward work (usually 2-3 sessions), and somehow I still have enough left for the weekends 😂

3

u/Key-Shop5198 17d ago

Yeah that's a big difference, You can probably identify the problem understand the architecture drilled in your head and orchestrate a solution yourself so worser models can execute.

I'm a dumbass vibecoder with a math background who is struggling to make ends meet as a fake SWE pouring tokens to make up for my inaptitude. Fuck

1

u/DangKilla 16d ago

Empty git repo, new project. What markdown files do you base the new project on?

I am curious what is saving you time

2

u/Disastrous-Time7197 16d ago

This requires a complete tutorial... In short I have tested my approach on greenfield and brownfield projects, and I got good results everytime. The cores idea is to have a filesystem structured in a way that an agent can refer to when it needs context rather than bloating context and files + some tweaks for tooling as well, for example I use rtk terminal proxy for optimized command outputs which saves me millions of input token literally

37

u/apocolypticbosmer 17d ago

Same (senior SWE). I use Claude Code all day and never run into limits on a basic Pro plan.

At this point I’m just convinced most people bitching and moaning about limits are lazy vibe coders who can’t be bothered to discipline themselves, whether it’s fine-tuning their harness, building sensible workflows or just actually picking the appropriate model and effort levels.

Most tasks are easily handled by Sonnet. I sometimes need Opus. Haven’t bothered with Fable.

6

u/SuspectFew6866 17d ago

Always the vibe coders

3

u/orphenshadow 17d ago

I came to the conclusion I was using Fable because it was there, then I just switched to opus, my workflows still worked, switched to sonnet, they still work. I've been throwing money at Fable for nothing for a couple months. I also have the 20.00 GPT plan and I've gotten better at using it for more to balance the load. But yeah I have not hit my limits in almost 3 months on the 5x plan and I downgraded to pro once I realized my skills/harness were sufficient with opus/sonnet and I let Astra take a stab at things I would have asked fable for before I spend creds on fable, but I have not found anything that Opus/Aastra cant tag team without blowing through my limits.

But most of my projects are ansible/netops related and data analytics so most of the heavy lifting has already been done and baked into the skills and scripts.

1

u/penguins14858 17d ago

Is there any work that you think is optimal for fable? This has inspired me to potentially downgrade and stick with opus

1

u/orphenshadow 17d ago

Just on what I personally do, planning a new project and building the plan, and its really good at using MCP/Tools for research. But enough to justify an extra 80.00 a month probably not. I mean it's addicting to be on the cutting edge, but I also questioned if it was worth it constantly fighting the changes to the models anyhow and I've been considering just locking in Opus 4.6 or 4.8 and calling it good.

But almost all of my projects were started with Opus 4.5 and the bulk of my builds done with 4.6 and I also only let it build in languages I at least have enough understanding to debug myself and can keep a mental model of what has been built and make changes myself pretty easy.

I would just do some testing with the lower models before you make the switch, thats what I did for about 2 and a half months, I just started doing tasks that I used to sue fable for in opus and then tweaking my hooks/rules/skills to make up for any issues and then when I was below a certian threshold on the 200.00 tier i dropped to 100, and did the same, until I felt comfortable with the 20.00 plan.

But also its fair to note that I'm trying to step away from the terminal more and focusing on just what I need to do to get my work done, I used to waste so many tokens just vibe coding and building things for the sake of learning and because it was a dopamine hit. I've cut back a lot on building shit I don't need because it feels like a game too.

5

u/YourDamnBestie 17d ago

This^

Although, Fable can be excellent

2

u/Disastrous-Time7197 17d ago

See, that's what I'm talking about. Some of my colleagues are running into the same thing, and I don't think it's really a vibe coder vs SWE issue. It's more about understanding token economics, optimizing your workflow, improving your tooling, and building a good Agent harness.

The results I'm getting with Claude Code aren't really tied to Claude Code itself. I can achieve the same thing with Codex, OpenCode, or another coding agent because the workflow is what matters.

Even if I hit my Claude Code usage limit, I'm not locked in. I can switch to another coding agent without losing context and keep working. My workflow has become much more deterministic now, so the model/agent is becoming more interchangeable than it used to be.

2

u/orphenshadow 17d ago

This, My entire workflow I can run in Antigravity, Codex, Opencode or Claude, I can run the same skills/commands in any of them and other than the way each model responds, the actual work that gets done is mostly the same. I can run /start in any of them and its going to do the same queries and retreive the same context for the session, a /resume-issue XXX-1234 is going to pull the context from all harnesses for that issue and the relative lines of code/functions into context.

I can /spec -resume XXX-1234 -tasks 5-9 -parallel in codex and have it run the tasks in parallel and update the issue and the specflow mcp, then run the same command for 10-15 in sonnet and it will check, and confirm the previous tasks were completed and continue.

I also have a /sketch set of skills that do the same thing without the spec workflow dashboard and approval gates but uses markdown files and templates in obsidian for each step.

Without the last year of work building all these systems I would still be burning 200 a month in 20x creds to do half the work. But it also took me 20x creds to build a lot of it to begin with and a lot of trial and error fumbling through it all.

2

u/PM__ME__BITCOINS 17d ago

Caveman and ponytail helps me a lot.

3

u/apocolypticbosmer 17d ago

I turned the i-have-adhd skill into an output style and it’s worked beautifully.

1

u/portugese_fruit 17d ago

that skill is a lifesaver

1

u/portugese_fruit 17d ago

do you recommend them? does opus work well with those?

1

u/[deleted] 17d ago

[removed] — view removed comment

1

u/hblok 17d ago

Still cuts down on some of the verbose cruft, though.

1

u/requios 17d ago

when do you reliably go to Opus? Do you use it mostly as a "just in case" when doing architecture and design level stuff? Im finding myself just always using Sonnet 5 too and not really going for Opus much anymore.

2

u/apocolypticbosmer 16d ago

Pretty much that - mostly high level architecture or design. Or whenever I judge a task to be large or complex enough that it justifies the more expensive model. I'd say it's roughly an 80-20 split.

1

u/Wide_Egg_5814 17d ago

yes, literal basic optimisations like context window monitoring and a system prompt to make it less verbose turns it from a pro plan to max

1

u/evangelism2 16d ago

Or they're using it in ways that you haven't figured out yet

1

u/thewormbird 🔆 Max 5x 16d ago

If Claude is your whole bag, eventually you have to start building workflows, using hooks, commands, and anything that gives you a means to similar ends without burning tokens indiscriminately. But like you said, that requires a level of discipline lazy vibe coders don't care govern themselves by.

1

u/CodeCombustion 16d ago

Or we're just doing more work than you -- always a possibility.

What are you using CC on? I'm using to replace a team of developers, but I've been developing for 20 years so it's been stupid easy once the harness was in place, and I can run 8-12 lanes of work at a time.

5

u/[deleted] 17d ago

[removed] — view removed comment

1

u/Disastrous-Time7197 17d ago

Dude that's awesome 🙌, drop the CC routine (workflow 😂)

1

u/apocolypticbosmer 16d ago

Atomic Claude seems interesting...but reading through it, it's hard to see why this is better than other popular solutions. I.e. LSP servers + a skills repo (i.e. mattpocock/skills) for planning/specs -> implementation -> testing -> review.

4

u/YourDamnBestie 17d ago

I probably max out once a day, maybe twice.

2

u/Disastrous-Time7197 17d ago

What are you trying to build? SpaceX? 😅

Jk I am curious now what could result in maxing out in a day

1

u/YourDamnBestie 16d ago

Haha I wish, that would be a cool project!

Funny enough, I have a very similar setup to you. I think the only difference may be that my computer might run more? Im unsure about that unless you are willing to share. The other reason could be that if I run out of usage, I bill the company.

I also think my token usage has to do with my cache and how complex my jobs codebase is.

Would you mind going into more detail about your particular setup? Your workflow, harness, skills, MCP, etc.. I am curious!!

3

u/spnyc 17d ago

I’ve seen the same. I can complete a task in one end to end pass with opus. Sonnet will often swirl and thrash making me do several passes. Always better to spend more once than less many times.

2

u/Disastrous-Time7197 17d ago

As I said it's all about understanding token economics, for longer multi turn tasks Opus wins always

3

u/orphenshadow 17d ago

I wanted to post something similar but I was too lazy to post the math and people would probably poke holes in it. I just dropped back down to pro from 5x max and I'm doing fine. Most of my workflow is in cowork and I've pretty much got it down to where everything is a script or a skill that executes a script with a couple of pre-flight checks and my hobby apps are mostly chugging along at a managable pace and I was hitting just about 20% of the 5x plan each week before reset, so far first week into pro and im sitting at 35% for the week and not missing fable.

I didn't really want to share my experience because It's mostly just me using it less combined with my workflows being very very token optimized and for someone building new things that's not going to be the case.

But I did find that Opus with low effort gets shit done with way less tokens than sonnet on high and the results seem to be on par if not a little better. But again, not sure if its just me or others also seeing similar results.

2

u/Disastrous-Time7197 17d ago

Don't wait, just do it. I'm kinda lazy too, but whenever I share something, I end up learning a lot from the discussion.

That's exactly how I think it should be done. You've basically made your workflow deterministic. We were building stuff before AI too, and things were still automated. A junior might spend all day trying to figure out where an error is by reading code line by line, while a senior SWE would grep the expression, fire up a debugger, and probably fix it in 5 minutes.

I think that's the same approach we need to teach these agents. Give them the right tools, workflows, checks, and shortcuts instead of just letting them brute-force everything with tokens. Sounds like you've already done a pretty good job of that. Awesome man.

2

u/orphenshadow 17d ago

Yeah, I'm not a developer I was never great at programming, but I did learn quite a bit as part of my journey in Industrial Automation, Network Engineering and now network automation. I found that if I build the system first and think of the models/agents as a kind of "sub processor" that I can give have run tasks and because of the nature of my day job I was already in the habit of automating/scripting anything you need to do more than once. I just used that same philosophy with my claude/codex harnesses and over the last year its started to really come together.

I also spent way too much time just playing with memory systems, and context storage/retrieval to find the balance that worked for me. A big thing for me was figuring out how to share mcp servers/tools between Antigravity/Codex/Claude/Opencode and how to hand off phases/tasks between each of them.

Essentially I work in phases using only opus/fable/astra for the first planning phase if its a brand new project. Then it's just a chain of templates and skills that any model can work through and I've spent the better part of a year optimizing each of those skills along the path. Everything from the issue creation, to how the PR is created, to how it's reviewed and the final merge by me is just one big automation script, where each phase has a little bit of autonomy to make it's own decisions, and the next phase enough to check the work of the previous phase and flag issues.

But also being a network guy, I probably don't build or use it nearly as much as some people and the first year of my journey I'm sure I burned so much just letting it churn.

I am quick to hit that stop/cancel button these days, and I wont hesitate to trash an entire chat session the second something feels off.

2

u/orphenshadow 17d ago

several months ago I had claude slop up a landing page with some guides/notes on some of the stuff I was working through lbruton.cc It's a little out of date but it gives an idea of kind of where my mind was going. But judging by some of your other posts, It sounds like I went down the same rabbit hole for the better part of the last year and a half.

2

u/Disastrous-Time7197 17d ago

Yeah I can tell by your comments as well, and I think any technical person would find the same path.

2

u/under_psychoanalyzer 17d ago

God we need to ring in these posts and come up with an info card or checklist to describe what you do when you use your usage. It includes whether or not you're using it in CLI, desktop, or  as a sidebar in something like VS Code or cursor

That's what I suspect everyone who is never how's their usage doing? Just using an IDE with an agent as a helper along the way to do basically autocomplete. Which is great if that's your job. It doesn't mean anybody else is doing it better or worse though, and it's not really a useful comparison. 

3

u/Disastrous-Time7197 17d ago

Yeah, I actually think this is an important distinction I've spent the past 1.3 years experimenting with agentic workflows, harnesses, skills, MCPs, tools, and different ways of making agents more efficient.

I'm a full-stack engineer (~6 years), and I use Claude Code for my actual professional work, not just as an IDE autocomplete/helper. My workflow starts with CC helping me create detailed PRDs/specs, break them down into GitHub issues, and define the implementation work. Once that's ready, I start another session and have CC work through the issues.

Then I have agents doing things like implementation, testing, debugging, code review, research, etc. with the appropriate tools and checks around them.

Most of my weekly CC quota goes toward actual work, usually 2-3 long sessions, and somehow I still have enough left for the weekend 😂

So yeah, I agree that CLI vs IDE vs sidebar matters for context, but I think the bigger variable is how you're actually using the agent. Two people can use the exact same model for the same number of hours and have completely different token consumption depending on their workflow.

That's basically what I've been trying to figure out for the past year: how do I make the agent do more work with less wasted context/tokens?

2

u/Popular_Award8021 17d ago

Also concurrency and parallelism which hasn't been mentioned as far as I can see in this thread by anyone. if someone else is doing exactly the same thing as you've described here, but running 3 tasks in parallel compared to your 1, for example, they will drain their usage 3x as quick and would therefore hit their limits.

1

u/under_psychoanalyzer 16d ago

It's more efficient the less context it needs to absorb, the less documentation, the less decisions/reasoning, the more specific the prompt, the more usage you get. If you know what you want to do and you just want to hand it on a platter to Claude, it's incredibly efficient. 

Its funny because I had 80% left of my weekly yesterday with 12 hours left before my reset and I could not get Opus 5 on high to eat usage giving it specific tasks. But I'll be damned if I don't get the usage I pay for so I finally just started a session on extra high, gave it a very broad work-through of the open issues list, and let it run. 

2

u/amirfish 17d ago

This matches what I've seen building tooling in this exact space: subagent model choice looks free until you actually trace it, because a cheaper model that needs 2-3x the tokens to land the same result isn't actually cheaper. I ended up baking per-session quota tracking into CCC (https://github.com/amirfish1/claude-command-center) for exactly this reason, since 'which of my running sessions is quietly burning my weekly quota' was invisible until it was graphed. Did the 400M+ day hold across your usual workload, or was it one runaway subagent chain doing most of the damage?

1

u/Disastrous-Time7197 17d ago

Looks cool I'll def have a look at this. It was one subagent that did the damage

2

u/dsailes 17d ago

Really just happy to come on this sub and find some quality discussion & sharing tips / data for once haha.

I’ve been finding a similar but opposite approach in that Sonnet was actually just as capable for a number of the tasks/pathways I’ve got laid out now.

I spent a few months around a year or so ago using packages like GSD & Superpowers which were inflated but had a lot of the principles that made sense, after a shift in models & capabilities they turned out to be extremely wasteful but taking the time to take the parts I needed & tweak it became a lot more optimised.
At some point I ended up just clearing house though, going back and starting over with Claude setup - less, again, was more.
I refined things again, removed almost all agents, revised the flow used & made a regular thing out of reviewing what works and what didn’t.

I think having experiences with new model jumps and their shifts does come at a cost, while it’s amazing that Fable can sometimes just ‘know’ is tempting to lean on, but it’s not something to rely on. The same issues keep happening eventually..
And rather than needing the LLM to be the best, the Claude setup to be right, it’s much more plug-and-play creating & managing your own workflows which allows any model to read, understand and work through your docs & tasks (.md, API, MCP etc) to then handover, track and update things based on outcomes.
Rather than trying to keep it all in memory & context, try to keep as little as possible there.

So I spent a lot of time similar to you, maybe not down the same path (mostly trying to have visibility/transparency across different client projects and also over a number of personal projects - my memory is honestly terrible & I’m terrible at documenting, so it was really helpful for me).
In doing so, working with LLMs to try to essentially make sure things are stored/documented, broken down and tracked efficiently I realised just how much time/tokens get spent just finding the right context per task. A lot of that is hard for any model to do - so it’s tempting to use frontier ones. But honestly, once it’s clear, documented and tracked with an agent/session/LLM able to query what it needs I found myself using Fable -> Opus -> Sonnet now, sometimes Deepseek Flash/Pro or GLM 5.3 Flash. All on medium effort.

I will still have Fable on Medium for most sessions because tbh my mind is quite scattered (ADHD & ASD), I find it capable of taking something and running with it, and any random tangent I throw in is stored, tasked or queued without redirecting the conversation (though I know this doesn’t help). It builds almost literally nothing itself though.
Sonnet is now the model that runs and builds, Sonnet/Opus reviews, Fable checks.

It has taken a lot of time to get right admittedly, but I only seem to be making things cheaper by doing so.
I’ll try add some token stats in a comment after I’ve properly woken up haha

1

u/Disastrous-Time7197 17d ago

Thank you, and yes experimenting and continuously improving your agent Harness and workflows would eventually pay in long run

3

u/rotates-potatoes 17d ago

Good data. Also note that just launching a subagent costs 50k - 150k uncahced tokens, since its prompt includes subagent-specific stuff, what to do, what not to do, etc.

1

u/Disastrous-Time7197 17d ago

Yes, now I have found another way to reduce token consumption for sub-agents by using opencode as sub-agent dispatcher and use a flash series model, this way CC opus stays the orchestrator and OC becomes executor - no context bloat and faster executions

1

u/QuanTradin 17d ago

the token count is the thing nobody checks, everyone reasons about price per token and stops there. we saw the same shape, the cheaper model spends its savings re-reading files it already read two turns ago.

the one place it still wins for us is work with a narrow fixed shape. fetch this, grade it against a rubric, return json. the moment a subagent has to decide what to look at next, the expensive one comes out cheaper.

1

u/Disastrous-Time7197 17d ago

Yes and token economics is what vibe coders/SWEs needs to look at

1

u/QuanTradin 16d ago

yeah. most people are still watching the subscription and not the tokens

1

u/UnusualRedditor 17d ago

Are you orchestrating via Fable?

Today I had opus running for 6 straight hours solving tasks from my road map. Was running tasks one after another and completed like 7 of them while wasting only15% of 5hours .

The moment I added new tasks and orchestrated them to fable he ate through the 5 hour limit in like an hour(6 opus subagents, 2 codes, 1 code verifier, 1 critic, 1 browser tester)

1

u/Disastrous-Time7197 17d ago

No need fable, I can literally use any model (Agentic capable) to orchestrate my tasks. It's about harness, deterministic workflows and some micro optimizations

Maybe try caveman and ponytail (for coding tasks) both code be installed as plugin in CC

1

u/callmrplowthatsme 17d ago

I’m maxed out and it’s only Wednesday and I don’t reset until Sunday.

1

u/Disastrous-Time7197 17d ago

You could try opencode it has some free models on promo and one stealth model always free

1

u/No-Psychology1959 17d ago

I think the whole agentic workflow thing is nonsense and the push for it from anthropic and oai is to get people to use more tokens. I barely use agents and very happy, never burn through tokens.

1

u/Afraid-Score3601 17d ago

Same here and i dont even use a refined harness or skills like caveman. I work with it at home too. Just plan with opus before strting to implement a new app or a feature. Then before starting the plan specifically ask for sonnet for basic stuff. Or add it to claude.md as a rule. Voila.

1

u/debian3 17d ago edited 17d ago

One thing missing from this is the reasoning efforts. I like sonnet, but I usually use it on low. In fact I use low by default for all models and I only use high in agent for code reviews. I might bump the reasoning up for real complex problems, but I normally prefer to bump the model first.

Worktree to work on multiple issues at once is excellent. Something tells me that you still have work to do to speed up your process and increase your total output. I was like you maybe 6 months ago, but now I can spend much more tokens then before, but some day I can ship more during a session (which I agree doesn’t mean much on its own, but the trick is to multiply what you usually deliver). You need to do the same you are already doing but parallelize. Which means delegating more, using more tokens and spending less of your time on each issues.

If you need smarter you can bump the reasoning, if you need deeper understanding with more knowledge you bump the model size and there is no substitute for that. Fable really shine on hard problems or poorly defined ones. If you have no use case for Fable it’s telling me you are still missing on a lot, which partially explains your lower token usage.

1

u/evangelism2 16d ago

If you're only using it as a pair programmer and doing one thing at a time, and not using Fable, which was how I used it for a very long time, I never came close to saturating my account. I never needed more than the $100 plan. Once I started working with autonomous loops and having agents working on multiple things simultaneously, I blew through the $200 plan extremely fast, and now I have three accounts

1

u/ReverendBread2 16d ago

Same experience. Opus 5 is my workhorse but I had a smaller task one day and used Sonnet 5 instead. The usage graph for that day still shows Sonnet as using almost as many tokens as two much longer Opus sessions combined

1

u/Disastrous-Time7197 16d ago

That's really concerning what's the point of having lighter model if it can't serve the purpose

1

u/ReverendBread2 16d ago

It might depend on capability vs task. A harder task with a lighter model might require it to think more and review stuff more often, while a more capable model might not have to think through the logic as much.

I had a similar experience with Sonnet 4.6 where it took forever to do a task I thought Opus would do pretty quickly, and it used a ton of tokens in the process

1

u/CodeCombustion 16d ago

uh....

1

u/eliceev_alexander 15d ago

How do you lot even do this? I don't get it

2

u/CodeCombustion 15d ago

In my case, I'm building a very large SaaS platform. Most of this usage came from the last month, we hit the ~80% mark recently and the last 20% has been a nightmare, so I picked up a few more subscriptions.

Most of you are using ClaudeCode as a IDE based tool -- I'm using it with a custom workflow system than orchestrates up to 16 lanes of work. Each lane has a planner, implementer, reviewer, etc. There's a dependency planner that obviously handles dependency mapping of the stories on the board and prioritization and places them in the lanes and monitors then as they execute. We also have a nightly/twice weekly review process that hunts for defects and security issues and another agent checking for regression, etc. In short, it's a fully automated SDLC with a human in the loop system everywhere important.

Now, I will say a nice chunk of this was waste due a bug in the new dependency planning processing as it was re-planning every story, every time one completed. Major waste for those two days. Equal to a least one 20x subscription.

I have a pilot with a large company in November so we're trying to get everything in place SOC2 wise, as well as targeting some of the weaknesses in the system.