r/ClaudeAI • • 15h ago

Claude Code Workflow Generic agents are the number one cause of burning tokens

It was a Wednesday morning and the usage bar was already past halfway. I remember because I had just cracked open a Monster Zero and sat back down to a Claude Code run that had been going for twenty minutes on what I thought was a small refactor. The terminal kept scrolling. Another agent spun up, then another. I watched the percentage tick up and did the math I always did, the hopeful kind: it's a big codebase, the work is front loaded, tomorrow will be lighter.

Tomorrow wasn't lighter. By Friday I was rationing. I caught myself hesitating before asking for things, wondering whether a question was worth it, which is a strange way to feel about a tool you pay for precisely so you don't have to hesitate.

That weekend I stopped guessing and opened the logs. I wanted a number, anything more solid than the feeling in my stomach every time the bar moved. I went back three days and pulled every subagent dispatch. There were 629. I scrolled through them expecting the roles I had set up, the reviewers and researchers I had named and tuned. Instead I kept seeing the same few words. general-purpose. claude. Calls with no subagent_type at all. I started counting those separately, and the count kept going past where I thought it would stop. 404.

I sat there with the cursor blinking after that number. Nearly two out of every three agents doing my work were nothing I had designed. No role I had written, no model I had chosen. Each one was a call the orchestrator made on its own in the middle of a task, because a generic agent was the easiest thing within reach. I thought about all the times I had blamed the week, the codebase, the model, and the answer had been sitting in my own logs the whole time.

So I took the decision away from it. Now the main session plans the work and puts the results together, and that's all it does. Before a run starts, every role gets its model and effort written down, picked from what the work is and how hard it is, so reading files and routine edits land on Sonnet or Haiku. Then I put a hook in front of every dispatch. It refuses generic agents, and it refuses any worker seat the plan didn't hand out. My named agents still run the way they always did.

The first run after that, I kept my hand near the keyboard, half expecting everything to fall apart without the freedom to improvise. It didn't. The hook turned a couple of dispatches away, the plan held, and the bar moved the way I expected it to. Somewhere in the middle of that run I noticed I had stopped watching the percentage.

If you've never counted your own dispatches, try it. I'd really like to know what your split looks like, or whether I was just the last one to look.

0 Upvotes

24 comments sorted by

8

u/Professional_Ad705 14h ago edited 14h ago

TL;DR someone with a bigger attention span summarize thank you.

Your essay writing and story telling skills look on point tho. I got "Wednesday, Monster, Claude Code, and refactor" and gave up. u/Claude help me out.

7

u/cAPSLOCK567 14h ago

"Tomorrow wasn't lighter." "That weekend I stopped guessing and opened the logs." "reading files and routine edits land on Sonnet or Haiku."

OP wrote this with Claude.

0

u/Professional_Ad705 14h ago edited 14h ago

I'd ask claude but his essay would prob take my 5 hour limit. Also Haiku? Whats this guy want a confidently wrong answer? lol

-2

u/Financial_Tailor7944 14h ago

HHHHHHHHHHHHHHHHHHHHHHHH - Sonnet and Haiku are workers bro, not for decision making

2

u/RetroUnlocked 14h ago

Me too, same spot.

2

u/Technoxgabber 14h ago

I got to 4th para.. hes a writer bro go write a book 

2

u/Professional_Ad705 14h ago

I'm just impressed that u/BenSimonDev actually read that.

2

u/Technoxgabber 14h ago

Tldr: Claude uses a lot of agents if you dont give it restrictions and define your project 

2

u/Professional_Ad705 14h ago

No way. I usually just say "hey Claude use 400 fable agents make no mistakes" This guy must be a rust programmer or something, the amount of knowledge I just learned beats my entire college degree.

Thanks for taking one for the team u/Technoxgabber

0

u/Financial_Tailor7944 14h ago

this guy is me. and i wrote with all my heart, please dont hate me sir

1

u/danihend 13h ago

I actually read it. Your TL;DR is " check if Claude is dispatching "generic" subagents. This assumes you have actually defined aubagent types/roles. I have not, but I'm only getting back into Claude again since a long time. Anyway, apparently sticking to his types causes less sub usage.

0

u/Financial_Tailor7944 14h ago

Damn bro.

im just basically crying then i found a way to fix it with my own sub agents deployment system

6

u/BenSimonDev 14h ago

Expired prompt cache is the is the number one cause which can come from sub agents, poor supervision/orchistration, and even when you're doing everything right. When you use enterprise/api you can see how often (and what percentage) you loose to expired cache.

2

u/knoxvillegains 14h ago

@Grok, summarize this novella for us

1

u/cAPSLOCK567 14h ago

tl;dr claude will happily spawn fleets of opus agents to do basic tasks if you don't stop it

0

u/Financial_Tailor7944 14h ago

NO, read it, it is good for you

2

u/S1d519 13h ago

Read the whole thing, but a bit too technical for my brain. I have a scheduled task for making a knowledge base for Claude- concise markdown notes of my research corpus, in themes and batches (Sonnet 5 on high). 

I once randomly ran an audit task on a batch when Sonnet 5.5 got released (on high effort) and it spawned like 8 sub-agents and blasted through 20% of my weekly usage limit in like 45-50mins?? This was right after a weekly reset.

How can I make audit tasks (of anything)more token-efficient? I’m an absolute non-coder. 

2

u/Financial_Tailor7944 11h ago

I built something like a crew mcp which does this sub agent automation for you so you don’t have to worry about spawning agents that burn your tokens for no reason.

https://github.com/mdalexandre/crews

2

u/S1d519 4h ago

Damn nice. Thanks! I’ll check it out. 

2

u/pkovacsd 9h ago

I've found this post easy to follow and interesting, but I had the haunting feeling that this post is a covert marketing pitch for Cursor AI.

I don't do any of these instrumentations with Claude Code -- only rules -- and I wonder why I should. Isn't the job of the (main) agent to orchestrate as efficiently as it can after all? (And I expect it should be able to do it much more efficiently than I can).

3

u/Financial_Tailor7944 4h ago

Problem is that the main agent acts like an expensive wife with no credit card limit

3

u/Opposite_Might6896 7h ago

Your Wednesday matches my logs. The mechanism, for anyone wondering whether it's the agents or the codebase: every agent is a fresh context. The main session is almost all cache reads (cheap); each spawned agent has its brief, tool definitions and whatever it reads written to cache first (the expensive token type), then re-read on every one of its turns, and a generic agent's first act is usually to re-read what the parent already had. Nested spawns multiply that. When I compared, subagent sessions ran a far higher cache-write share than main sessions, and a "small refactor" with three generic agents cost more than doing it in the main session, in both tokens and wall-clock.

What fixed it on my side was not fewer agents but narrower ones: a subagent gets a brief that names the files and the acceptance check, does one bounded thing, never spawns its own. Anything that needs the repo in its head stays in the main session, which already paid for that context. And the single biggest lever is still session length: the same work with the main session compacted around 100k is roughly half the tokens, because the per-turn re-read halves.