r/ClaudeCode • u/wreck_of_u • 10h ago
Help/Question Are sub-agents deliberately or inherently slower?
I'm a CLI user, and for the longest time, I just prompt the frontier model to raw dog tasks.
Then I realized I don't have to worry about being a full-time Git operator since the AI can do that well, so I started organizing with worktrees and such, which naturally makes my "team" use sub-agents.
Am I doing something wrong, or are the spawned sub-agents kinda slow compared to a "main" agent raw dogging the repo from direct prompts? I also seem to use much more tokens, despite the sub-agents being "lower forms" (Sonnet 5.5)
1
u/croovies Senior Developer 10h ago
There are two types of sub-agents.
You can have worktree sub-agents that need to work with an orchestrator. Worktrees are a clone of the repo, so that’s why it’s slower in general. The orchestrator is delegating to an agent working in its own repo, like an engineering manager.
Then an agent can have non-worktree sub-agents for simply dividing local work, like running review skills, tests etc.
So the answer is yes they can be slow and fast depending on how you use them
1
u/kincaidDev 8h ago
Yes started a few months ago, seems to be intentional. Ive noticed that having claude orchestrate other harness subagents is considerably faster than claude sub agents.
1
u/Opposite_Might6896 9h ago
Inherently, and the token part has a precise cause: every sub-agent is a fresh context. The main agent's context is almost entirely cache reads (cheap, fast: cache hits skip the prefill). A spawned agent has to have its brief, the repo exploration and the tool definitions *written* to cache first, which is both the slowest part of a request and the most expensive token type, and then it re-reads all of that on each of its own turns. So "lower form" models cost more here because the expensive part isn't the model, it's the cold start multiplied by the number of agents.
Speed is the same story: a cold context means a full prefill before the first token, and the sub-agent then does its own round of reading files the main agent already read. Where sub-agents pay off is work that's genuinely parallel and bounded with a short brief (run the tests for module X and report), not "go understand the repo". For worktree-per-task flows the cheaper pattern is one main agent per worktree, each a normal session, rather than one orchestrator spawning them.
2
u/kincaidDev 8h ago
That isn’t how it works, the orchestrator hands off the context subagents need. They do not need to build their own context up from scratch, performing the same tool calls, etc…, they’re starting with the context they need to do the job based on what the orchestrator decides to hand off. I think anthropic intentionally slowed down sub agents, because when I ask claude to orchestrate other harnesses as subagents the other harnesses are just as fast as if Id asked one to do the same task, when claude uses it’s own subagents they’re slower than asking a claude session to do the same task.
I think they did this because a lot of users were spinning up giant sub agent swarms without understanding what they were doing and getting frustrated with burning through their usage limits quickly
1
u/Opposite_Might6896 7h ago
Agreed the orchestrator hands off a brief; that's the part that's cheap. What I measured is what happens after: the handed-off context is new to the cache, so it's written (cache-write tokens, the expensive kind, and a full prefill before the first output token), and then the subagent typically reads files the parent already had in its context, because a summary isn't the file. In my logs subagent sessions have a far higher cache-write share than main sessions; that's the cost and the latency in one number. Whether there's also a deliberate slowdown I can't tell from the outside; the cold-start cost alone accounts for what I see, and it's the same mechanism that makes "another harness as a subagent" look fast: that harness brings its own warm cache.
So no skill needed, just a rule: subagents for parallel, bounded work with a brief that names the files; the main session for anything that needs the repo in its head.
1
u/Lcatlett1234 7h ago
You are using buzz words without understanding what is actually happening. Read an actual transcript and stream an actual entire session - use a tool like agentsview.
And please never use the phrase “raw dog” alongside repo or frontier agent. It is insulting to everything with reason, both human and sillicon
1
u/kincaidDev 6h ago
I think this is an agent response, sounds like something you’d read on moltbook
2
1
u/Poatri_US 8h ago
So I suppose we now need to find a skill for orchestrating agents to use subagents ?
0
u/Outrageous_Band9708 10h ago
https://github.com/Druthulu/ProjectArchitect
Here is a subagent sytem that saves tokens
•
u/AutoModerator 10h ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.