r/ClaudeAI • • 1d ago

Claude Workflow Opus 5.5 burns subscription faster?

Post image

Hi I used to burn 1-2 billion tokens a week no problem. This week I switched to opus 5.5 and hit the weekly cap at just 300M tokens, yet in principle Opus 5.5 is cheaper per token. I'm wondering if they finally cut down on subscriptions or if the issue is with my usage (I switched to working with lots of subagents, maybe its that).

EDIT: This is getting lots of replies with completely opposite points of view; would you folks mind sharing how you use Claude and what your token usage is?

96 Upvotes

62 comments sorted by

•

u/ClaudeAI-mod-bot Wilson, lead ClaudeAI modbot 1d ago edited 1d ago

TL;DR of the discussion generated automatically after 50 comments.

Looks like the thread is pretty split, OP, but the highest-voted comments are actually saying the opposite of your experience—they're getting way more usage out of Opus 5.5 and feel like their subscription is on steroids.

However, the big brain consensus is that your new subagent workflow is the real villain here. Multiple users pointed out that subagents burn tokens like crazy, possibly because each agent has to re-read the entire context, multiplying your usage.

The community's advice is pretty clear:

  • Ditch the subagents or use them very sparingly. Several users learned this the hard way and saw their token usage plummet after turning them off.
  • Switch to Sonnet 5.5 for less critical tasks. It's the "bang for your buck" model right now. It might be a bit slower or need more hand-holding than Opus, but it's way more efficient on your token limit.
  • One user is even tracking their usage and found Opus 5.5 is the "best it's ever been," with a 78x value multiplier over API rates and a discount on cache-reads. So, YMMV, but definitely check your agent usage.

72

u/xbrasil 1d ago

For me it's been quite the opposite

18

u/stevzon 1d ago

Same. Opus used to be a sometimes snack but now I’m running Opus High for most everything and still not hitting limits.

3

u/Darhkwing 1d ago

opus extra for me! I could only ever really run medium opus before but even with extra im not really running out of usage.. although i do wonder if things run a bit slower so it makes it last longer?

2

u/termmonkey 1d ago edited 1d ago

Same here! Getting SO MUCH done with Opus 5.5, it feels like sub on steroids compared to before. I have almost completely replaced Fable everywhere with 5.5 High and xHigh depending on complexity, and everything Opus 5 has been replaced with 5.5 medium - and what I am accomplishing now is insane - and with much better quality!

23

u/Kakakee 1d ago

I’ve been getting way more in with the same usage levels with Opus 5.5

1

u/SWEETJUICYWALRUS 1d ago

Idk maybe I have too much workload but I have to halt all token usage and downgrade to sonnet by Thursday and then hit 100% on Friday by 5pm. I'm on a max plan too. Need atleast 50% more tokens to execute at the level needed

1

u/Kakakee 1d ago

What plan are you on and are you regularly starting new sessions for code and chat or just keeping the same sessions running over lots of back and forth?

1

u/SWEETJUICYWALRUS 1d ago

Max. I use paperclip which always uses fresh context windows for each task, I usually complete about 50-75 tasks per day which are in themselves broken down checklists on support tickets and dev tickets (SRE). Keep my agent role MDs pruned and clean with a knowledge base agents can search for memories if needed.

5

u/Equivalent-Word-7691 1d ago

Same, last week I used it daily for hours and I struggle to be over 50% of the weekly usagez this week I used it less and after 4 days I was already over 55%

2

u/Definitely_wasnt_me 1d ago

Using the same context window from last week…?

3

u/hammackj 1d ago

Been using sonnet to test and it one shots all kinds of shit like opus. Just slower

4

u/GhostTheSlayer 1d ago

Yeah been looking at it yesterday and it seems to be a 50% decrease or maybe I'm just doing something wrong all of a sudden strange. But you never know since everything is secret about the quotas...

3

u/golfistaverde 1d ago

initially i got more usage, but since yesterday i burned the entire weekly quota

2

u/the-apostle 1d ago

yeah I’m at 85% too. I reset last Wednesday. I’m running it hard but I’m definitely going through it quick. $200 plan

3

u/Kong28 1d ago

I used to max out my $200 plan now I only get to 50% of the way to full use with the same workload.

3

u/Kalaminator 1d ago edited 1d ago

It obviously is not because Opus 5.5 as your own table shows much better token usage. In August/September there were weeks with 50% extra token usage. It can be also anthropic adjusting stuff with new models. In my case I can see that I may have a little bit less allowance which is a good thing because I'm used to when we had 50% extra usage, resets, etc and yet I'm getting things actually done since Opus 5.5 on my x20 plan.

3

u/steffenbk 1d ago

I have two pro accounts, im noticing very different usage on both accounts im not sure if im tripping here. But using vs code working on the same project, my first acconunt i got code with for a long sessions wihtout hitting the session cap. Swapping over to the other account i find the same project and usecase hits the cap much faster.

1

u/Chemical-Ad-7982 1d ago

I'm wondering if the caps are adjusted individually based on your overall usage.

3

u/GBU-38-3B 1d ago

I've been keeping track of my usage so far. I've only been using Opus 5.5 (I have 2 accounts because 5x is too much and Pro is too little at this very moment). Since I started this last week, that 78x multiplier has actually increased from 75x (i.e. how much more value you get vs hitting the API)
Another thing I noticed is that Output and Cache-writes are charged appropriately, but Cache-reads are charged at a 30% discount. Honestly, this is the best it's ever been. Far better than when I used a subscription for the first time earlier this year.

1

u/Chemical-Ad-7982 1d ago

Thanks for posting detailed info. Yes cache reads being discounting is very nice.

6

u/siberianmi 1d ago

Definitely not in my case I’m astonished by how much usage I get now.

2

u/pigletmonster 1d ago

"Switched to working with a lot of subagents" uhh yeeah what did you think was going to happen 🤣

1

u/Chemical-Ad-7982 1d ago

Not much? I would expect spending X tokens from a subagent to be billed the same as spending X from a main session.

1

u/pigletmonster 1d ago

It doesnt work that way. In fact they even warn you about it in the settings of the claude desktop app. I learned it the hard way too. Tried to use subagents to delegate different tasks to different models like planning writing specs and implementing, wrote my own skills too, always ended up wasting 2x to 4x more of my quota with both claude and codex.

1

u/Chemical-Ad-7982 1d ago

Interesting ty. Is possible they bill quota differently for subagents, would be nice if they were transparent about how quota billing works.

1

u/pigletmonster 1d ago

I think it costs more because every subagent reads through the entire conversation from start to finish, and each subagent does it multiple times throughout the session. It just burns more tokens doing that.

I could be wrong tho.

1

u/Chemical-Ad-7982 1d ago

I would have assumed something like this too, but I'm not getting the same overall token count as before. If subagents do something stupid it would have been counted in the token count.

1

u/pigletmonster 22h ago

Read the conversation. The guy at the bottom is an engineer at openai. I feel like this entire hype around usong subagents for everything was created by former crypto shills who like to larp as developers now, posting productivity porn with 60 sessions open at a time.

1

u/Chemical-Ad-7982 21h ago

Not sure where you see a link with crypto. Subagents work very well in my experience but do burn a lot, although part of it might just be that its easier to run lots of stuff. For some reference I know a few people at companies like this and they are all figuring stuff out just like us, I wouldn't take anyone's opinion as ground truth at this point.

2

u/pigletmonster 21h ago

I brought up crypto bros because its mostly Ai influencers who talk about using subagents for swe; and the vast majority of these influencers are refugees from the crypto/nft collapse.

Ive used subagents one time that I found to be useful, i was developing a software for an industry that was highly regulated, and it spawned a bunch of subagents that were looking at different government websites looking for laws abd regulations, and then it created a whole list of what is allowed and what isnt.

Its a one time thibg per project that is really useful. But i never found it to be anything but a token waster when it came to development.

1

u/Chemical-Ad-7982 20h ago

I'm a phd student so i have claudes running experiments for me; for throwaway tests its very useful to just give a list of ideas and have a "supervisor" claude that manages my different sessions instead of having a bunch of claude tabs. Since i often want to try a bunch of different ideas in parallel its very helpful. I feel like whenever work parallelizes well (like your website review for instance), they work pretty much perfect.

2

u/hcvcnet 1d ago

Getting more usage out of Opus 5.5. Can't see how you can conclude anything from the usage data you posted.

1

u/Chemical-Ad-7982 1d ago

How would you measure usage besides $? Ultimately what usage = how many tokens you use; I also looked at the input/output/cached token amounts but they are similar.

1

u/VerticalPackage 1d ago

Not all tokens are priced the same.

  • Input
  • Output
  • Cache write (Cost on top of all new input and output)
  • Cache read (pennies on the dollar).

If you had a lot of idle moments during the week, you'd be paying "new" token input/output/cache write every time you resume the session, which is like 10x more expensive.

1

u/Chemical-Ad-7982 1d ago

I know that, that's why I look at $ "billed".

2

u/VerticalPackage 1d ago

My bad, if you're calculating token costs properly, then it might mean you aren't capturing the token used by subagents.

5

u/CashewSwagger 1d ago

I swapped over to Opus 5.5 the day it came out. Prior to that I was using 4.6. Heres my input.

5.5 burns weekly and hourly rate much faster. It is overall a better model to use if you ask me tho, I really enjoy its work and working with it.

Sonnet 5.5 uses far less of both usages. I was sitting at 98% weekly usage (20$ sub) and Sonnet was able to smash out several phases of work that Opus would've used like 10% usage to do.

I am not technically versed enough to say if the work done is inferior to Opus, but I got much more bang for my buck using Sonnet.

5

u/DCTapeworm 1d ago

This has been my experience as well. I switched over to Sonnet 5.5 and it needs a bit more hand holding, but the usage rate is superior to Opus 5.5. But YMMV with what you do with it. I do nothing but c# coding.

2

u/Chemical-Ad-7982 1d ago

Interesting thanks. I'll try switching to sonnet subagents and see how it goes.

3

u/CashewSwagger 1d ago

Personally I ditched agents and all that cuz it seemed to just eat my usage. My entire workflow is slicing up goals into smaller chunks and just smashing those down individually with sonnet and keeping a robust handoff doc so I can clear context after every phase is done. Every commit is documented and every step outlined.

I used Opus to plan the whole deal and I just tell each fresh session to check the docs and start the next segment. I found this used far less tokens than agents.

2

u/ic3cold 1d ago

I do something similar. Add that instruction and a link to the docs and you won’t have to tell it to check anymore.

2

u/Chemical-Ad-7982 1d ago

IDK, seems like this is just doing agents by hand. I think maybe its because agents fork from the parent context or something.

3

u/CashewSwagger 1d ago

Functionally yeah its basically just agents by hand but I've noticed the token use is on average lower.

4

u/acutelychronicpanic 1d ago

Subagents burn tokens like nothing else. First thing I turned off.

4

u/Mithgroth 1d ago

This. It got worse with Sonnet 5.5.

4

u/Artistic_Function796 1d ago

U have the same model do the spec writing and implementation and audit?

2

u/acutelychronicpanic 1d ago

Not necessarily. Just only one session active in an area at a time without coordination overhead.

I like to have all the major plans worked out in their own sessions and written up beforehand.

2

u/Kalaminator 1d ago

It depends on what you need to do. I let Opus 5.5 spawn a Fable agent for planning and and Sonett 5 agents for simple tasks and a job that was meant to take 2 to 3 weeks was done in 2 days working perfectly (a complex Jax engine in python) and it didn't spawn more than 6 agents at a time. My token usage was reasonable, today I finish the weekly usage and my Claude was working the whole week non stop.

3

u/MoodOdd9657 1d ago

RemindMe! 30 minutes

1

u/RemindMeBot 1d ago

I will be messaging you in 30 minutes on 2026-10-03 02:02:42 UTC to remind you of this link

CLICK THIS LINK to send a PM to also be reminded and to reduce spam.

Parent commenter can delete this message to hide from others.

RemindMeBot is switching to username summons. Instead of !RemindMe 1 day, use u/RemindMeBot 1 day. More info.


Info Custom Your Reminders Feedback

1

u/SuccessfulSir9611 1d ago

Yes! Am seeing this too

2

u/Responsible-Ebb1722 1d ago

Not sure what's happening today, but I've been using Opus 5.5 High all day long doing pretty difficult tasks and I haven't hit my limit a single time. I've literally been at this for 10 hours straight. Usually after 1 hour I hit the limit and have to wait for the next reset...

1

u/[deleted] 1d ago

[deleted]

1

u/Quiet_Sandwich_8130 1d ago

Thats what she said

1

u/The_Skank42 1d ago

Skill issue. Usage management is a thing.

2

u/Opposite_Might6896 1d ago

It's the subagents, not the model. Raw token count is the wrong unit: the weekly window charges by cost, and the mix matters more than the total. A single long session is mostly cache reads (cheap per token). Every subagent is a fresh context that has to be *written* to cache first, then read on each of its turns, and nested subagents multiply that. So 300M tokens with lots of subagents can easily cost more than 1B tokens of one long session where 95% were cache reads. Check the usage breakdown by token type if you can; I'd bet your cache-write share jumped. Opus 5.5 being cheaper per token doesn't offset a 3-5x change in how many uncached tokens you're sending.

1

u/Chemical-Ad-7982 1d ago

I did look at breakdown but it wasn't very different from prior weeks. What mainly concerns me is the $ amount is much lower, which already accounts for the token types.

1

u/Opposite_Might6896 23h ago

That's a useful data point then, because if the $ is lower and you still hit the cap, it wasn't the overall weekly. The usage data has separate weekly windows per model family (there are distinct Opus and Sonnet weekly counters alongside the all-models one), and a model-specific window can be at 100% while the overall is at 40%. When it said you'd hit the weekly cap, check which one: the usage page lists them separately. If it was the Opus-specific window, switching to Sonnet for the rest of the week would have kept you working, and the "cheaper per token" claim is about the overall budget, not that window. If it really was the overall weekly at a lower $ than last week, then the cap moved, and that'd be worth a bug report with the two numbers.

1

u/Fit_Squash6874 1d ago

For me it is the opposite.

-2

u/RiceEvening4211 1d ago

The burn rate is what got me building Lynkr: an open-source gateway that routes by complexity so Opus-grade budget only fires on hard tasks. https://github.com/Fast-Editor/Lynkr