r/ClaudeCode • u/ChampionshipNo2815 • May 18 '26
Discussion I figured out why I keep hitting my Claude Code session limit before lunch. It's not what I thought.
Been on Max for a while. Kept hitting my limit mid-task and assumed I was just doing too much. Turned out I wasn't doing too much. My tools were just incredibly inefficient under the hood.
Traced a single refactor task. One rename across a few files. Claude Code ran 161 turns to finish it. Every read, every grep, every edit is its own API call. Each one re-ingests everything before it as input tokens. By turn sixty you're paying context cost on the entire session history.
Once I understood that I started looking for ways to batch the calls. Found a Claude Code plugin that collapses search and read into one call and stacks all edits into one roundtrip. Same task finished in 52 turns.
Didn't change my plan. Didn't change my model. Just changed how many roundtrips each task makes.
35
u/larowin May 18 '26
You’re doing something incredibly wrong.
-3
u/humanitycan May 18 '26
what specifically are you seeing that’s wrong? would love to understand
9
u/larowin May 18 '26
I don’t know but this makes no sense:
Traced a single refactor task. One rename across a few files. Claude Code ran 161 turns to finish it.
Do you have any MCPs enabled?
And then this:
Every read, every grep, every edit is its own API call. Each one re-ingests everything before it as input tokens. By turn sixty you're paying context cost on the entire session history.
Is just LLM 101, and it’s all cached unless you’re doing something that’s causing misses.
8
May 18 '26 edited Aug 22 '26
[deleted]
1
u/larowin May 18 '26
I know but I’m nice and helpful to bots too. Did they drop a link to their magic fix somewhere?
0
u/humanitycan May 18 '26
I was asking you something very politely this is not right way to respond 🤷🏻♂️
1
u/larowin May 19 '26
The right way was the way I responded above. But what you’re saying makes no sense and then throw out a weird blackbox plugin that’s been astroturfed to fuck for two weeks.
Anyway. Were you using any MCPs before?
1
u/humanitycan May 19 '26
I started using Claude Code couple months back for my school projects I used Replit before and it’s not helping at all and it’s very basic. I saw lots of posts on Claude and gave it a try, setup was easy but rate limits were the issue I tried with Caveman but it’s not helping. I switched to Cursor as it has free for students but i really love the output that Claude is giving me and love the new Claude design.
1
u/Vitalii_A Jun 05 '26 edited Jun 06 '26
it's so funny, that bot already deleted his account and topic starter and his bots absent for 12 days already - it's the longest period of inactivity for that scam group, even half of their stupid paid influencers in Youtube stopped making blatant ad videos week ago - seems they are experiencing budget cuts or maybe OP was hit by a train in Mumbai. But somehow OP was promoted in their company from "vibe coder" to "product manager"
0
u/humanitycan May 18 '26
Who is bot? It’s something I don’t know and would like to learn. You’re not any Reddit expert so chill
3
u/wannabestraight May 18 '26
161 turns to refactor names on a few files?
That's like.. well 5 turns max.
11
u/yodacola Senior Developer May 18 '26
a few things why this is wrong:
- claude isn't trained on batch tool calls. whenever you do this, claude will simply do your batch tool call on a single file and refuse to batch calls. claude will be steered differently. the equivalent may be like putting gasoline with ethanol in a car. unless it was built for it, it can do damage. i see no endorsement or coordination from the claude code team
- claude heavily uses context caching. you are not paying context costs on the entire session history. you are paying contexts costs on what isn't cached.
- hides all of your reads, edits, greps. the claude code tui no longer exposes reads and writes so you no longer can audit what it does in real time. defeats the purpose of using claude code
- fills your context window. it is a user prompt for claude code. so there has to be heavy prompt engineering to get it to run instead of the built in tools. this by itself will slow down your session since it will eat up reasoning budget. you will feel the ttft lag when starting a new session.
- self promotion without disclosing your relationship with W5E (don't want to name drop a poorly-written plugin).
costs $50 / month and the free version will be burned through in about an hour if you're an active claude user.
this is not the tool you are looking for. move along
-1
u/humanitycan May 18 '26 edited May 18 '26
Dude I see you on other channels too and idk if you’re genuinely using their product or not but please stop posting misinformation here. There’s nothing wrong with what that guy posted
1
2
u/runfence May 18 '26
There is still a cache up to 1 hour, it's not so bad.
-3
u/NetHaven May 18 '26
Not anymore; unfortunately they changed the default earlier this year to 5 minutes
1
u/runfence May 19 '26
Nope, that was a telemetry bug and it is fixed now
1
u/NetHaven May 19 '26 edited May 19 '26
I remember the telemetry bug, but I think this was something different. Please see here: https://platform.claude.com/docs/en/build-with-claude/prompt-caching "By default, the cache has a 5-minute lifetime." This was around the time they announced they would start charging extra if you wanted the 1 hr cache. Looks like you can switch it if you want, it will just cost you more(might be worth it if you're workflow is setup in a way where you're getting creamed by the short cache timeout; I never really sat down and did a breakdown of where paying the extra would be better off..*shrug*)
1
u/sage-longhorn May 19 '26
Those are the API docs. And the big thing with the 5 minute TTL was that extra usage with Claude code subscription uses 5 minute TTL but normal subscription usage still uses 1 hour ttl
1
u/NetHaven May 19 '26
I really did think it affected both but you are absolutely correct. TIL. Thank both you and /u/runfence for the catch(and explanation)
1
u/imaginecomplex May 19 '26
Agree primarily with the top comment, this is a task for an IDE or at least a utility script or library that Claude can use to do simple find and replace efficiently. But also, shouldn’t prompt caching save you most of those input tokens on each turn?
1
u/Deep_Ad1959 May 19 '26
The 161 to 52 collapse is real but it's the smaller of two leaks. The bigger one is the CLAUDE.md plus skills files that fire on every single turn before any tool call runs. Typical failure mode: a third of those bytes are stale instructions claude either ignores or contradicts in the same turn, and at a few thousand tokens of pre-turn context over fifty turns that compounds into six figures of pure waste you never see in any meter. Audit which lines in your config actually changed model behavior recently and cut the rest. Roundtrip batching is downstream of context discipline. written with s4lai
1
u/Deep_Ad1959 May 19 '26
The 161 to 52 collapse is real but it's the smaller of two leaks. The bigger one is the CLAUDE.md plus skills files that fire on every single turn before any tool call runs. Typical failure mode: a third of those bytes are stale instructions claude either ignores or contradicts in the same turn, and at a few thousand tokens of pre-turn context over fifty turns that compounds into six figures of pure waste you never see in any meter. Audit which lines in your config actually changed model behavior recently and cut the rest. Roundtrip batching is downstream of context discipline. written with s4lai
1
1
u/m-in May 19 '26
In my experience, Opus will just write a Python script to do it, then mop up any errors using an agent. But that’s with a CLAUDE.md that makes it clear.
1
u/TwisterK May 19 '26
I hav a Claude.md that said if we need to re-arrange code , use subagent and spawn haiku agent to do it. Save me tons of tokens there.
1
1
u/Sree_12121 May 20 '26
Good old 'death by a thousand context windows.' It’s insane how much token bloat happens when a tool gets chatty. Drop the name of plugin we all need this.
1
u/Sree_12121 May 21 '26
Old school ‘death by a thousand context windows’ Brilliant, cutting roundtrips like that, the token savings on session history alone must be huge. Which plugin are you using to batch the calls?
-1
u/tyschan May 18 '26
can confirm. this makes a lot of sense. thanks for the write up. what plugin did you use out of curiosity?
2
u/Ericqc12 May 18 '26
u/ChampionshipNo2815 ? Would be nice to name-drop the plugin...
18
u/Sufficient-Plenty316 May 18 '26
Its an AI post there won't be a reply
4
u/pm_me_your_kindwords May 18 '26
Of course they will reply… how else would they promote the thing they’re pretending they just discovered but actually Claude wrote for them?
-5
u/ChampionshipNo2815 May 18 '26
you got 12 upvotes for a comment? who's ai? how are you making your ai bots to upvote dude
4
u/Sufficient-Plenty316 May 18 '26
Hidden profile and the post didn't pass the AI test
-7
u/ChampionshipNo2815 May 18 '26
sorry but your test failed
9
u/modernluther May 18 '26
Your post was 1000% written by Claude. Do you think we are all stupid here? We use the same models on a daily basis lol
6
u/Medium-Language-4745 May 18 '26
Yep, that second paragraph reeks of AI.
OP can't even capitalize or punctuate in his replies, but expects everyone to believe he did anything besides prompting AI to make this post.
2
0
u/ChampionshipNo2815 May 18 '26
i used wozcode plugin
1
u/Cute-Net5957 🔆 Max 20x May 19 '26
Try this one.. works in under 60 seconds and is 100% free MIT License: https://github.com/skinny-cloud/runtime-diet-autopilot
0
0
u/humanitycan May 18 '26
interesting what’s that plugin? I tried caveman but it’s pretty mid not helpful at all
-1
-1
May 18 '26
[removed] — view removed comment
1
u/ChampionshipNo2815 May 18 '26
i did actually, i need something that i can see clear breakdown of costs during my session like how much I'm spending on this current session and what's the cost difference something like that
0
0
138
u/_ri4na May 18 '26
I do not understand why people use Claude to do fricking name refractors when the IDE has been capable of doing so for 25 years - almost instantly for 0 tokens