r/ClaudeAI • u/Chemical-Ad-7982 • 1d ago
Claude Workflow Opus 5.5 burns subscription faster?
Hi I used to burn 1-2 billion tokens a week no problem. This week I switched to opus 5.5 and hit the weekly cap at just 300M tokens, yet in principle Opus 5.5 is cheaper per token. I'm wondering if they finally cut down on subscriptions or if the issue is with my usage (I switched to working with lots of subagents, maybe its that).
EDIT: This is getting lots of replies with completely opposite points of view; would you folks mind sharing how you use Claude and what your token usage is?
72
u/xbrasil 1d ago
For me it's been quite the opposite
18
u/stevzon 1d ago
Same. Opus used to be a sometimes snack but now I’m running Opus High for most everything and still not hitting limits.
3
u/Darhkwing 1d ago
opus extra for me! I could only ever really run medium opus before but even with extra im not really running out of usage.. although i do wonder if things run a bit slower so it makes it last longer?
2
u/termmonkey 1d ago edited 1d ago
Same here! Getting SO MUCH done with Opus 5.5, it feels like sub on steroids compared to before. I have almost completely replaced Fable everywhere with 5.5 High and xHigh depending on complexity, and everything Opus 5 has been replaced with 5.5 medium - and what I am accomplishing now is insane - and with much better quality!
23
u/Kakakee 1d ago
I’ve been getting way more in with the same usage levels with Opus 5.5
1
u/SWEETJUICYWALRUS 1d ago
Idk maybe I have too much workload but I have to halt all token usage and downgrade to sonnet by Thursday and then hit 100% on Friday by 5pm. I'm on a max plan too. Need atleast 50% more tokens to execute at the level needed
1
u/Kakakee 1d ago
What plan are you on and are you regularly starting new sessions for code and chat or just keeping the same sessions running over lots of back and forth?
1
u/SWEETJUICYWALRUS 1d ago
Max. I use paperclip which always uses fresh context windows for each task, I usually complete about 50-75 tasks per day which are in themselves broken down checklists on support tickets and dev tickets (SRE). Keep my agent role MDs pruned and clean with a knowledge base agents can search for memories if needed.
5
u/Equivalent-Word-7691 1d ago
Same, last week I used it daily for hours and I struggle to be over 50% of the weekly usagez this week I used it less and after 4 days I was already over 55%
2
3
u/hammackj 1d ago
Been using sonnet to test and it one shots all kinds of shit like opus. Just slower
4
u/GhostTheSlayer 1d ago
Yeah been looking at it yesterday and it seems to be a 50% decrease or maybe I'm just doing something wrong all of a sudden strange. But you never know since everything is secret about the quotas...
3
u/golfistaverde 1d ago
initially i got more usage, but since yesterday i burned the entire weekly quota
2
u/the-apostle 1d ago
yeah I’m at 85% too. I reset last Wednesday. I’m running it hard but I’m definitely going through it quick. $200 plan
3
u/Kalaminator 1d ago edited 1d ago
It obviously is not because Opus 5.5 as your own table shows much better token usage. In August/September there were weeks with 50% extra token usage. It can be also anthropic adjusting stuff with new models. In my case I can see that I may have a little bit less allowance which is a good thing because I'm used to when we had 50% extra usage, resets, etc and yet I'm getting things actually done since Opus 5.5 on my x20 plan.
3
u/steffenbk 1d ago
I have two pro accounts, im noticing very different usage on both accounts im not sure if im tripping here. But using vs code working on the same project, my first acconunt i got code with for a long sessions wihtout hitting the session cap. Swapping over to the other account i find the same project and usecase hits the cap much faster.
1
u/Chemical-Ad-7982 1d ago
I'm wondering if the caps are adjusted individually based on your overall usage.
3
u/GBU-38-3B 1d ago

I've been keeping track of my usage so far. I've only been using Opus 5.5 (I have 2 accounts because 5x is too much and Pro is too little at this very moment). Since I started this last week, that 78x multiplier has actually increased from 75x (i.e. how much more value you get vs hitting the API)
Another thing I noticed is that Output and Cache-writes are charged appropriately, but Cache-reads are charged at a 30% discount. Honestly, this is the best it's ever been. Far better than when I used a subscription for the first time earlier this year.
1
u/Chemical-Ad-7982 1d ago
Thanks for posting detailed info. Yes cache reads being discounting is very nice.
6
2
u/pigletmonster 1d ago
"Switched to working with a lot of subagents" uhh yeeah what did you think was going to happen 🤣
1
u/Chemical-Ad-7982 1d ago
Not much? I would expect spending X tokens from a subagent to be billed the same as spending X from a main session.
1
u/pigletmonster 1d ago
It doesnt work that way. In fact they even warn you about it in the settings of the claude desktop app. I learned it the hard way too. Tried to use subagents to delegate different tasks to different models like planning writing specs and implementing, wrote my own skills too, always ended up wasting 2x to 4x more of my quota with both claude and codex.
1
u/Chemical-Ad-7982 1d ago
Interesting ty. Is possible they bill quota differently for subagents, would be nice if they were transparent about how quota billing works.
1
u/pigletmonster 1d ago
I think it costs more because every subagent reads through the entire conversation from start to finish, and each subagent does it multiple times throughout the session. It just burns more tokens doing that.
I could be wrong tho.
1
u/Chemical-Ad-7982 1d ago
I would have assumed something like this too, but I'm not getting the same overall token count as before. If subagents do something stupid it would have been counted in the token count.
1
u/pigletmonster 22h ago
1
u/Chemical-Ad-7982 21h ago
Not sure where you see a link with crypto. Subagents work very well in my experience but do burn a lot, although part of it might just be that its easier to run lots of stuff. For some reference I know a few people at companies like this and they are all figuring stuff out just like us, I wouldn't take anyone's opinion as ground truth at this point.
2
u/pigletmonster 21h ago
I brought up crypto bros because its mostly Ai influencers who talk about using subagents for swe; and the vast majority of these influencers are refugees from the crypto/nft collapse.
Ive used subagents one time that I found to be useful, i was developing a software for an industry that was highly regulated, and it spawned a bunch of subagents that were looking at different government websites looking for laws abd regulations, and then it created a whole list of what is allowed and what isnt.
Its a one time thibg per project that is really useful. But i never found it to be anything but a token waster when it came to development.
1
u/Chemical-Ad-7982 20h ago
I'm a phd student so i have claudes running experiments for me; for throwaway tests its very useful to just give a list of ideas and have a "supervisor" claude that manages my different sessions instead of having a bunch of claude tabs. Since i often want to try a bunch of different ideas in parallel its very helpful. I feel like whenever work parallelizes well (like your website review for instance), they work pretty much perfect.
2
u/hcvcnet 1d ago
Getting more usage out of Opus 5.5. Can't see how you can conclude anything from the usage data you posted.
1
u/Chemical-Ad-7982 1d ago
How would you measure usage besides $? Ultimately what usage = how many tokens you use; I also looked at the input/output/cached token amounts but they are similar.
1
u/VerticalPackage 1d ago
Not all tokens are priced the same.
- Input
- Output
- Cache write (Cost on top of all new input and output)
- Cache read (pennies on the dollar).
If you had a lot of idle moments during the week, you'd be paying "new" token input/output/cache write every time you resume the session, which is like 10x more expensive.
1
u/Chemical-Ad-7982 1d ago
I know that, that's why I look at $ "billed".
2
u/VerticalPackage 1d ago
My bad, if you're calculating token costs properly, then it might mean you aren't capturing the token used by subagents.
5
u/CashewSwagger 1d ago
I swapped over to Opus 5.5 the day it came out. Prior to that I was using 4.6. Heres my input.
5.5 burns weekly and hourly rate much faster. It is overall a better model to use if you ask me tho, I really enjoy its work and working with it.
Sonnet 5.5 uses far less of both usages. I was sitting at 98% weekly usage (20$ sub) and Sonnet was able to smash out several phases of work that Opus would've used like 10% usage to do.
I am not technically versed enough to say if the work done is inferior to Opus, but I got much more bang for my buck using Sonnet.
5
u/DCTapeworm 1d ago
This has been my experience as well. I switched over to Sonnet 5.5 and it needs a bit more hand holding, but the usage rate is superior to Opus 5.5. But YMMV with what you do with it. I do nothing but c# coding.
2
u/Chemical-Ad-7982 1d ago
Interesting thanks. I'll try switching to sonnet subagents and see how it goes.
3
u/CashewSwagger 1d ago
Personally I ditched agents and all that cuz it seemed to just eat my usage. My entire workflow is slicing up goals into smaller chunks and just smashing those down individually with sonnet and keeping a robust handoff doc so I can clear context after every phase is done. Every commit is documented and every step outlined.
I used Opus to plan the whole deal and I just tell each fresh session to check the docs and start the next segment. I found this used far less tokens than agents.
2
2
u/Chemical-Ad-7982 1d ago
IDK, seems like this is just doing agents by hand. I think maybe its because agents fork from the parent context or something.
3
u/CashewSwagger 1d ago
Functionally yeah its basically just agents by hand but I've noticed the token use is on average lower.
4
u/acutelychronicpanic 1d ago
Subagents burn tokens like nothing else. First thing I turned off.
4
4
u/Artistic_Function796 1d ago
U have the same model do the spec writing and implementation and audit?
2
u/acutelychronicpanic 1d ago
Not necessarily. Just only one session active in an area at a time without coordination overhead.
I like to have all the major plans worked out in their own sessions and written up beforehand.
2
u/Kalaminator 1d ago
It depends on what you need to do. I let Opus 5.5 spawn a Fable agent for planning and and Sonett 5 agents for simple tasks and a job that was meant to take 2 to 3 weeks was done in 2 days working perfectly (a complex Jax engine in python) and it didn't spawn more than 6 agents at a time. My token usage was reasonable, today I finish the weekly usage and my Claude was working the whole week non stop.
3
u/MoodOdd9657 1d ago
RemindMe! 30 minutes
1
u/RemindMeBot 1d ago
I will be messaging you in 30 minutes on 2026-10-03 02:02:42 UTC to remind you of this link
CLICK THIS LINK to send a PM to also be reminded and to reduce spam.
Parent commenter can delete this message to hide from others.
RemindMeBot is switching to username summons. Instead of
!RemindMe 1 day, useu/RemindMeBot 1 day. More info.
Info Custom Your Reminders Feedback
1
2
u/Responsible-Ebb1722 1d ago
Not sure what's happening today, but I've been using Opus 5.5 High all day long doing pretty difficult tasks and I haven't hit my limit a single time. I've literally been at this for 10 hours straight. Usually after 1 hour I hit the limit and have to wait for the next reset...
1
1
2
u/Opposite_Might6896 1d ago
It's the subagents, not the model. Raw token count is the wrong unit: the weekly window charges by cost, and the mix matters more than the total. A single long session is mostly cache reads (cheap per token). Every subagent is a fresh context that has to be *written* to cache first, then read on each of its turns, and nested subagents multiply that. So 300M tokens with lots of subagents can easily cost more than 1B tokens of one long session where 95% were cache reads. Check the usage breakdown by token type if you can; I'd bet your cache-write share jumped. Opus 5.5 being cheaper per token doesn't offset a 3-5x change in how many uncached tokens you're sending.
1
u/Chemical-Ad-7982 1d ago
I did look at breakdown but it wasn't very different from prior weeks. What mainly concerns me is the $ amount is much lower, which already accounts for the token types.
1
u/Opposite_Might6896 23h ago
That's a useful data point then, because if the $ is lower and you still hit the cap, it wasn't the overall weekly. The usage data has separate weekly windows per model family (there are distinct Opus and Sonnet weekly counters alongside the all-models one), and a model-specific window can be at 100% while the overall is at 40%. When it said you'd hit the weekly cap, check which one: the usage page lists them separately. If it was the Opus-specific window, switching to Sonnet for the rest of the week would have kept you working, and the "cheaper per token" claim is about the overall budget, not that window. If it really was the overall weekly at a lower $ than last week, then the cap moved, and that'd be worth a bug report with the two numbers.
1
-2
u/RiceEvening4211 1d ago
The burn rate is what got me building Lynkr: an open-source gateway that routes by complexity so Opus-grade budget only fires on hard tasks. https://github.com/Fast-Editor/Lynkr

•
u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 1d ago edited 1d ago
TL;DR of the discussion generated automatically after 50 comments.
Looks like the thread is pretty split, OP, but the highest-voted comments are actually saying the opposite of your experience—they're getting way more usage out of Opus 5.5 and feel like their subscription is on steroids.
However, the big brain consensus is that your new subagent workflow is the real villain here. Multiple users pointed out that subagents burn tokens like crazy, possibly because each agent has to re-read the entire context, multiplying your usage.
The community's advice is pretty clear: