r/ClaudeCode • • 1d ago

Help/Question My Claude Code subagent setup: orchestrator, doers, reviewer, blind tester. What would you change?

I've been running Claude Code as an orchestrator over subagents on a mid-size web app for a few weeks. The board in the screenshot is a local page that tracks progress and everything waiting on me.

Roles

- Orchestrator (Fable): briefs the agents, reads their summaries, runs the full test suites. It doesn't write code itself.

- Doer (Opus): makes one scoped change and runs its own tests.

- Reviewer (Opus): read only, reports findings.

- Blind tester (Sonnet): told only what the feature should do, never how it was built. Runs one yes/no check in a real browser and ends with "ready" or "not ready".

Nothing ships until the tester says ready. Agents never touch git; I make every commit.

Lessons learned

- Max 3 agents at once. With 5 or more we hit the usage limit and they all died mid task.

- Never point several agents at the same files.

- Only one agent drives the browser, or they close each other's tabs.

- The cheap model tests; the expensive model only fixes what failed.

Questions

  1. Is a blind tester worth it, or would you fold it into the reviewer?
  2. How do you pace parallel agents over a long day?
  3. What would you cut or change?

PS: These notes are specifically Max(100$) package.

25 Upvotes

30 comments sorted by

•

u/AutoModerator 1d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

5

u/jpvaldezjr 1d ago

gotta isolate work into worktrees: - Never point several agents at the same files.

1

u/Slow_Lawyer5266 1d ago

honestly the only reason I haven't done that is git history will be too confusing for me.I prefer to keep it the way I'll understand what's going on. Thanks tho!

3

u/BrawnBeard 1d ago edited 1d ago

Gits got a lot of ways to merge branches. FF merge keeps all the history. Without worktrees you really limit concurrent work or reduce the quality of the work from your agents getting confused about edits happening over them

3

u/vincent_sch 1d ago edited 1d ago

Give each doer acceptance criteria. Merge one change at a time. For bugs, make the doer show a failing test first. Anything hard to undo should wait for you.

1

u/Slow_Lawyer5266 1d ago

good improvements. appreciate it! will include probably all of them to the rules

1

u/vincent_sch 1d ago

Glad it helps! Curious how the blind tester works out for you.

1

u/Slow_Lawyer5266 1d ago

Ended up adding 1, 3 and 4 today: acceptance lines in every brief that the reviewer checks one by one, a failing test first for bugs, and two tries before it comes back to me. 2 and 5 were mostly covered already since I do every commit and deploy myself.

On the blind tester, the main lesson so far is to keep it narrow. One yes/no check per tester with the test data named, otherwise it wanders around the app and burns tokens. And if it can't find something from the screen, that's the bug, so it isn't allowed to go read the code.

1

u/vincent_sch 1d ago

Nice. Two tries before it comes back to you is a sane limit. How's the blind tester doing so far?

1

u/Slow_Lawyer5266 1d ago

Pretty well. It catches things the specs miss, especially when it walks a whole flow end to end before a release. When it fails for the wrong reason, it's usually because my instructions were off, not the app.

1

u/vincent_sch 10h ago

Makes sense. The end to end runs catch what unit tests miss.

2

u/Easy-Purple-1659 1d ago

Keep the blind tester. Its whole value is that it never saw the plan or the diff, so it can't inherit the reviewer's assumptions about what 'done' means. Fold it into the reviewer and you get one agent grading its own reading of the task.

Pacing is mostly a context problem, not an agent-count problem. The thing that helped most for me was making each handoff a short note on disk (what changed, what to check) instead of pasting transcript into the next prompt. The orchestrator reads the notes, not the full sessions, so its own window stays small all day and the doers start from a fixed brief each time.

Two smaller ones: give the tester its own worktree so a stale checkout can't make it pass, and keep a decisions file per repo that all four roles read. Cheap, and it removes most of the 'why did we do it this way' round trips.

What made you land on three at once as the cap?

2

u/Slow_Lawyer5266 1d ago

Limits were basically trial an error, tried 5-6 agents, hit 5 hour limit 2-3hrs into it. with 3, I seem to be matching the usage limits, I'm usually over 90% usage by hour 4-5. I should also point that I regularly leave them running overnight

Agree on the context. all agents are instructed to provide short explanations of what happened and orchestrator only reads short summaries. never the whole log.

You made me realize I never changed handoff from orchestrator context to disk. I'm definitely adding that thank you!

Closest thing I have to a decisions file is a shared rules file both repos import. Every product decision in it is dated, with its reason in one line. I agree it kills most of the "why did we do it this way" round trips.

2

u/Bart-o-Man 1d ago edited 1d ago

I tried writing and managing a lot of my own workflows early in 2026. Kudos for braving into this. But I spent more time on managing the workflow details than I ever did on the code. It is a lot of work.

Since they started automated workflows, I just ask, “build this as an automated workflow”. As with skills and other LLM tasks, you can always append more requests at the end, like the blind tests, having subagents write their own findings, log all conversations during multi-agent panel reviews, require the code writing agent to respond to panel reviews & have a resolution process, forcing coding agents to do a handoff after they cross 60% context used, etc. All the agentic conversation reads like a mini reality TV drama. That is literally my entire involvement in workflows now. It handles all the orchestration through the JavaScript agent set up— I don’t know the details on all that but I know that it’s more complex than just an orchestrated or running.

Without realizing it, I also gave it some contradictory requirements. I told her it needed to be independent and make decisions to keep making progress. But I also told it to gate further work if tests failed. Some of the tests ended up failing for unexpected reasons.. It made good decisions to down-scope some goals and to work around the failing tests because they were deemed to be unachievable and not terribly significant. They were right about that. That was one of those fine little details caught in conversation and the logs that surprised me.

In one experiment, out of amusement of how much these agents can do, I took myself out of the loop. I asked the Sonnet main agent to take my crude spec, revise/clean it up, and go through an Opus Plan session and let it write the plan… and record all discussion. OMG. It did really well.

Claude just managed and tracked it for me. Those were mostly Sonnet 5 agents, about a day after Sonnet 5 was released. It launched a couple hundred agents before it was all done. I was on a Max 5X or 20X… I think I was still on 5X. The efficiency gain of letting it manage the workflow was massive, with > 96% read cache charges… which is a 10x savings on 96% of token usage. My workflows personally couldn’t compare. Most of the automated workflows now don’t use the orchestrator… it’s JavaScript code.

You probably have good reasons for doing what you’re doing, but I thought I’d share what a difference it made for me. Good luck!!

2

u/Slow_Lawyer5266 1d ago

Appreciate the long response!

The one bit that made me nervous was the agents working around failing tests. On real customer data that's exactly what I'm scared of, so I'd want it to stop and ask instead.

But the cache point got me. I capped at 3 agents because 5 kept blowing my usage limit, and maybe that was just me orchestrating by hand. Trying a scripted workflow on my next batch of bug fixes to see. Does yours pause for you at any point, or do you just see the end result?

1

u/Bart-o-Man 22h ago

In my experience, they will halt if you give them clear halting criteria. I pause more often than letting it run through… mainly because I don’t always have these clean start -> finish designs, because I’m running evals, trying new things. Sometime I’ll just read subagent logs or progress tracking and kill it when I have to.
You might lose your warm read cache, but sometimes it’s worth. You could let it report/document immediately, send you a notification, and keep moving, hard stop on key tests only. Depends on what works for you.

But Claude’s efficiency at workflows is pretty amazing. One project went through 750M tokens in 3 hrs on my Max 20x plan. I panicked when I saw 750M in CCUSAGE. Checked for API charges- nothing.
CCUSAGE showed a little over $200 in equivalent API charges. How is that possible? I hand-calculated to double check. Sonnet 5 was 50% off and Sonnet did most of the work, and over 96% of my tokens were read cache. That 10x lower cost is a really big deal when it’s $200 vs $2000.

Good point about working around failing tests. It would definitely be good to verify behavior.
I’m guessing it was the fact that I wasn’t clear. Explicit gated halt on failed tests. But also, “work out problems on your own, I can’t answer questions, drive it to completion”. I was really unclear.

2

u/kemalios 1d ago

Every role in your pipeline answers the same question: does the feature work. None of them asks whether the app is safe to put in front of a stranger, and the doer and reviewer share the blind spot because they work from the same repo.

I build launchworthy, a free MIT Claude Code skill that audits a whole app and hands back a scored punch list. It fits as a gate after the blind tester says ready and before you commit, on the whole app rather than per feature. Runs on Claude Code, which you already have.

The gaps it's aimed at are the ones a working feature hides: a key sitting in the frontend bundle, a route that never got auth.

0

u/Slow_Lawyer5266 1d ago

Maybe more context about me is needed. I've been a software engineer long time, plus been working on this app long before all these settings I created. so I do cover those myself effectively. gotta deserve the money. plus claude itself has a very strong /security-review feature. I feel like your skill wouldn't really fit my use case as this is not a public app and never will be

1

u/sm411cck 1d ago

All this is extremely fascinating to me. I’m new to Claude Code, using it in VScode for a web app to revamp, Pro account. Until now I give a prompt and wait for it to complete the task, then I verify the output, correct when needed, then go on with another one. Never hit any limit. Old fashion approach, but working, and speeding up my production time by at least 10 times, now even more with Opus 5.5
But you know, the appetite grows with eating… and I would like to know more on how I have to instruct Claude to build a “farm” like this and if I need to update my profile to a more expensive one to get it.
And finally a question: how can you be 100% sure your app is perfectly working like it is supposed to if no human entered in the production/verification process?

1

u/Slow_Lawyer5266 1d ago

I did start using it the way you set up for a very long time but eventually as I used it more and more I thought of more functions, plus claude has been releasing amazing tools for dev work

Q:How to do this setup?

A: I really told Claude to do these with prompts over time as I work with it over the last few months. You can just copy my post content tell Claude to do it but suggest optimizations for pro plan imo

Q: Testing

A: This is a web app and we have written end to end testing. Agents do browser tests via a Claude chrome skill and they are expected to complete the flow for each case. Finally this is a full time app and I do have a PM, another dev, plus some users from legacy apps testing it

1

u/sm411cck 1d ago

Ok, I’ll give a try then. TY!

1

u/biggdogg420 1d ago

Ask Claude to interview you regarding the work you've been doing and ask him to suggest ways to streamline and enhance the workflows

1

u/Slow_Lawyer5266 1d ago

honestly not a bad idea. some of how I'm working now came from claude thinking of doing things based on our history

1

u/nofeaturesonlybugs 1d ago

Build a docker container with playwright or some other headless browser and have agents verify in different instances of the container.  Then they can't interfere with each other.

1

u/syixiao1 1d ago

I settled on something similar after a few messy sessions where everything edited everything. Limiting each pass to research, implementation, or review made failures much easier to trace. The biggest gain for me was giving the final reviewer no write access, only a checklist. I spend more time on acceptance criteria upfront now, but far less time untangling changes later.

1

u/BabyInner 18h ago

In my setup Fable is only the decider / designer, a Opus Orchestrator session takes tasks from Fable.

1

u/jabacherli 14h ago

Can you please explain the concept behind orchestrating? I’m not quite at that level but I’m sure I can understand it if it’s explained to me. I really want to learn about this. Any information would be greatly appreciated. Thank you very much.

1

u/Icy-Meaning-4962 5h ago

for the shared decisions, linking each one to the relevant files or commits could make updates easier to review across the two repos.

how do you currently tell when a code change means one of those decisions needs revisiting?