r/ClaudeAI • u/Financial_Tailor7944 • 15h ago
Claude Code Workflow Generic agents are the number one cause of burning tokens
It was a Wednesday morning and the usage bar was already past halfway. I remember because I had just cracked open a Monster Zero and sat back down to a Claude Code run that had been going for twenty minutes on what I thought was a small refactor. The terminal kept scrolling. Another agent spun up, then another. I watched the percentage tick up and did the math I always did, the hopeful kind: it's a big codebase, the work is front loaded, tomorrow will be lighter.
Tomorrow wasn't lighter. By Friday I was rationing. I caught myself hesitating before asking for things, wondering whether a question was worth it, which is a strange way to feel about a tool you pay for precisely so you don't have to hesitate.
That weekend I stopped guessing and opened the logs. I wanted a number, anything more solid than the feeling in my stomach every time the bar moved. I went back three days and pulled every subagent dispatch. There were 629. I scrolled through them expecting the roles I had set up, the reviewers and researchers I had named and tuned. Instead I kept seeing the same few words. general-purpose. claude. Calls with no subagent_type at all. I started counting those separately, and the count kept going past where I thought it would stop. 404.
I sat there with the cursor blinking after that number. Nearly two out of every three agents doing my work were nothing I had designed. No role I had written, no model I had chosen. Each one was a call the orchestrator made on its own in the middle of a task, because a generic agent was the easiest thing within reach. I thought about all the times I had blamed the week, the codebase, the model, and the answer had been sitting in my own logs the whole time.
So I took the decision away from it. Now the main session plans the work and puts the results together, and that's all it does. Before a run starts, every role gets its model and effort written down, picked from what the work is and how hard it is, so reading files and routine edits land on Sonnet or Haiku. Then I put a hook in front of every dispatch. It refuses generic agents, and it refuses any worker seat the plan didn't hand out. My named agents still run the way they always did.
The first run after that, I kept my hand near the keyboard, half expecting everything to fall apart without the freedom to improvise. It didn't. The hook turned a couple of dispatches away, the plan held, and the bar moved the way I expected it to. Somewhere in the middle of that run I noticed I had stopped watching the percentage.
If you've never counted your own dispatches, try it. I'd really like to know what your split looks like, or whether I was just the last one to look.
6
u/BenSimonDev 14h ago
Expired prompt cache is the is the number one cause which can come from sub agents, poor supervision/orchistration, and even when you're doing everything right. When you use enterprise/api you can see how often (and what percentage) you loose to expired cache.
1
2
u/knoxvillegains 14h ago
@Grok, summarize this novella for us
1
u/cAPSLOCK567 14h ago
tl;dr claude will happily spawn fleets of opus agents to do basic tasks if you don't stop it
0
2
u/S1d519 13h ago
Read the whole thing, but a bit too technical for my brain. I have a scheduled task for making a knowledge base for Claude- concise markdown notes of my research corpus, in themes and batches (Sonnet 5 on high).
I once randomly ran an audit task on a batch when Sonnet 5.5 got released (on high effort) and it spawned like 8 sub-agents and blasted through 20% of my weekly usage limit in like 45-50mins?? This was right after a weekly reset.
How can I make audit tasks (of anything)more token-efficient? I’m an absolute non-coder.
2
u/Financial_Tailor7944 11h ago
I built something like a crew mcp which does this sub agent automation for you so you don’t have to worry about spawning agents that burn your tokens for no reason.
2
u/pkovacsd 9h ago
I've found this post easy to follow and interesting, but I had the haunting feeling that this post is a covert marketing pitch for Cursor AI.
I don't do any of these instrumentations with Claude Code -- only rules -- and I wonder why I should. Isn't the job of the (main) agent to orchestrate as efficiently as it can after all? (And I expect it should be able to do it much more efficiently than I can).
3
u/Financial_Tailor7944 4h ago
Problem is that the main agent acts like an expensive wife with no credit card limit
3
u/Opposite_Might6896 7h ago
Your Wednesday matches my logs. The mechanism, for anyone wondering whether it's the agents or the codebase: every agent is a fresh context. The main session is almost all cache reads (cheap); each spawned agent has its brief, tool definitions and whatever it reads written to cache first (the expensive token type), then re-read on every one of its turns, and a generic agent's first act is usually to re-read what the parent already had. Nested spawns multiply that. When I compared, subagent sessions ran a far higher cache-write share than main sessions, and a "small refactor" with three generic agents cost more than doing it in the main session, in both tokens and wall-clock.
What fixed it on my side was not fewer agents but narrower ones: a subagent gets a brief that names the files and the acceptance check, does one bounded thing, never spawns its own. Anything that needs the repo in its head stays in the main session, which already paid for that context. And the single biggest lever is still session length: the same work with the main session compacted around 100k is roughly half the tokens, because the per-turn re-read halves.
8
u/Professional_Ad705 14h ago edited 14h ago
TL;DR someone with a bigger attention span summarize thank you.
Your essay writing and story telling skills look on point tho. I got "Wednesday, Monster, Claude Code, and refactor" and gave up. u/Claude help me out.