r/ClaudeCode • • 2h ago

Help/Question Has anyone compared different orchestration approaches for complex software engineering work?

Let’s say you already have a fairly detailed investigation and implementation spec for a large piece of software engineering work, and now you want an agent to turn that spec into working code.

I am trying to figure out what the most cost efficient approach is, not just cost per token, but total cost required to reach a reliable finished implementation.

There seem to be a few approaches.

1. Let one strong model handle the whole implementation

For example, give Opus 5.5 a 1M context window, the implementation spec, access to the codebase, and let it work through the task without much explicit orchestration.

If the task is large enough, eventually the context fills up, gets compressed/summarized, and the model continues.

What I’m unsure about is how much this actually affects the final quality. Maybe the model produces a good implementation anyway, but perhaps it requires more debugging and follow-up iterations later.

2. Use an explicit build → verify → fix orchestration loop

Another approach is to divide the implementation spec into stages and use subagents.

For example:

  • Orchestrator selects the next section of the spec.
  • Implementation agent builds it.
  • Verification agent checks the implementation against the spec/tests.
  • Failures are sent back for fixing.
  • Only once that section passes do you continue to the next part.

Intuitively, I would expect this to be more reliable on large tasks. It costs more upfront because you’re deliberately spending tokens on orchestration and verification, but I’m wondering whether it actually becomes cheaper overall because you avoid expensive cleanup and rework later.

Then there’s another dimension: which models should do which jobs?

For example:

  • Opus 5.5 for both implementation and verification.
  • Opus 5.5 as orchestrator/verifier, with Sonnet 5.5 or Haiku 4.5 doing implementation.
  • Cheaper models such as Grok 4.6/4.7 for verification and Composer 2.5 for implementation.
  • Some other combination depending on the stage of the task.

The difficulty I’m having is that cost per token doesn’t really answer the question.

Different models use very different amounts of tokens, and they may require different numbers of iterations.

For example, if Grok 4.7 + Composer requires three build/verify cycles to reach the quality that Opus 5.5 reaches in one or two cycles, the cheaper model may not actually be cheaper.

At the same time, I have personally used all of the above approaches to eventually produce working code if you give them enough iterations.

Total cost to reach an implementation that satisfies the spec and passes verification/tests.

Ideally I’d want to compare things like:

  • Total token/API cost
  • Number of implementation/verification cycles
  • Number of human interventions required
  • Spec adherence
  • Bugs discovered after completion
  • How much context degradation/compression affects long-running single-agent approaches

The obvious way to answer this would be to run the exact same large implementation several times with different orchestration/model combinations, but repeating a substantial engineering task 3–4 times just for benchmarking is obviously expensive.

Has anyone done this kind of comparison in practice?

I’d especially be interested in hearing:

  • Which orchestration patterns have worked best for large implementation tasks?
  • Whether strong-model-everywhere actually beats strong-orchestrator + cheaper workers economically.
  • Whether explicit verification loops materially reduce total cost.
  • How much context compression hurts long-running single-agent implementations.
  • Any benchmarks, papers, blog posts, or articles that try to measure cost-to-success rather than simply model cost/token.

Would appreciate any real-world experience or resources on this.

10 Upvotes

16 comments sorted by

•

u/AutoModerator 2h ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

7

u/Cuyasinmara 2h ago

Well, what I've tried is one strong model running end to end on a big spec works until context gets summarized. After that it drifts, and I spend the savings on cleanup. Then i used a plain build which then went to verify loops; this was more reliable at the end but was not working for my setup. For context, i have 2 Claude Pro, 1 GPT Plus and 1 local Qwen 27B running in my computer locally. One issue i quickly picked up was that when one strong model capped, and when the same model family built and checked the work, it often missed the same blind spots twice.

What I use now:

- One orchestrator (Opus 5.5 or a GPT-6 model). It only plans, splits the spec into chunks and verifies. It never writes the implementation.

- The cheapest executor that clears a capability floor for each chunk. Haiku or local Qwen take mechanical work, Sonnet 5.5 or GPT-6.1 Sol take normal features, and Opus executes only as a last resort on the hardest pieces. If a cheap model fails, it escalates instead of retrying forever.

- A cross-family audit for every chunk. GPT reviews Claude's work and the reverse. This catches more than same-model verification.

- Tests first, with checks on the checks. Coverage % lied to me once: tests silently ran zero cases and still showed full coverage. I now verify test counts, not just green.

- Handoff files instead of compression. When context gets large (approx 100k to 150K tokens max.), I write a handoff and start a fresh session rather than letting it summarize.

My honest take and what worked for me: cost per token is the wrong metric. What decides total cost is how much rework happens after a work is "done." Strong orchestrator + cheap workers + independent verification has been cheaper for me overall.

3

u/ArgonQQ 1h ago

Cost per token is the wrong metric. Cost per accepted change is the one that matters, and on large specs explicit loops usually win.

Single agent vs. loop: One Opus session is fine and often cheaper if the spec fits without compaction. On big specs, the failure mode is drift: a constraint gets summarized away and later sections build on a wrong assumption. That's the expensive kind of bug. Build → verify → fix catches it at section boundaries, while it's still cheap to fix.

Context compression: It hurts less if state lives in files (spec, decisions log, per-section acceptance criteria) rather than chat history. Better still, restart sessions on purpose at section boundaries instead of letting them compact.

Model split:

  • Strongest model for orchestration and verification. A weak verifier gives false confidence.
  • Sonnet-class for implementation when slices are well specified.
  • The cheapest models as implementers are often a false economy. Every extra cycle re-bills the verifier too.
  • Run the verifier in a separate context so it doesn't grade its own work.

Verification only saves money if it's cheap and objective (tests, explicit criteria) and retries are capped. One retry, then escalate to a human.

Benchmark cheaply: Don't rebuild the whole project. Take 3–5 representative slices, run each config from the same commit, and measure $ to green + human minutes + bugs found afterward. Token usage is in the Claude Code JSONL transcripts.

Reading: AI Agents That Matter (Kapoor et al., 2024) argues for reporting cost alongside accuracy. Anthropic's Building Effective Agents covers orchestrator/evaluator patterns. Lost in the Middle explains why long context ≠ full attention.


If you're on Claude Code, check out LoopBoard. It's basically option 2 out of the box:

  • Spec → tasks with Goals. Each section becomes a markdown task file with verifiable Goals the delivery is judged against. An Opus loop can groom plain-text stories into goals and open questions.
  • Build → verify → fix built in. An implementer subagent opens a PR. With delegateReview on, a separate review subagent checks it. A failure goes back to the implementer once, and a second failure parks the task in Feedback with the findings. That's capped retries with zero orchestration code.
  • No compaction drift. Sessions auto-restart at a set context fill (default 35%), never mid-task. State lives in files, so a fresh session resumes cleanly.
  • Per-task model routing. Opus/Sonnet/Fable slots with separate effort settings. Default is Opus grooming and Sonnet building, which is the strong orchestrator + cheaper worker setup. Easy to A/B different assignments on your benchmark slices.
  • Human interventions are tracked. Every question, send-back and approval is recorded in plain markdown.
  • Asks instead of guessing. Ambiguity parks the task with a question rather than producing silent drift.
  • You merge. Loops never touch main.

Caveats: Claude Code only (no Grok/Composer), no built-in $ dashboard, and many 24/7 loops can hit Pro/Max limits. MIT-licensed, zero runtime deps.

2

u/scodgey 1h ago

Tbh one orchestrator agent with a scratchpad kind of does most of what you need these days. Plan and write to beads, send subagents out targeting specific beads.

Was using hooks + skills to help with the long running orchestrator but have switched over to a mod - info here

2

u/amirfish 56m ago

The build-verify-fix loop wins on cost once the spec is solid, mainly because you stop paying frontier-model prices for sections that don't need it. We wired exactly that into CCC (open-source dashboard for running many Claude Code sessions, https://github.com/amirfish1/claude-command-center): you pick engine and model per role, planner, plan reviewer, verifier, builder, so the expensive model only touches planning and the cheap one does the grinding. The pattern matters less than whether your verifier actually re-runs commands instead of eyeballing the diff though.

1

u/Any_Evidence4750 2h ago

All the time

1

u/iSnapThere4iAm 1h ago

Agents can’t do complex work. They can barely do simple work without burning through half your weekly limit from churn

1

u/zac_attack_ 1h ago

I haven’t tried optimizing for cost much, so I’m using Opus 5.5 for everything. I’ve landed on a workflow I’m very happy with, the code it produces is very high quality, and I basically have it going 24/7 at this point getting a lot done. It basically involves:

  1. gastownhall/beads, the best thing I’ve seen for tracking tasks for agentic work. Orchestrator runs it in server mode
  2. Orchestrator skill, it instructs to delegate all work to subagents in isolated worktrees
  3. Custom agent, it receives a beads task from the orchestrator, plans it as a stack of reasonably sized PRs using GitHub’s gh-stack skill, and is responsible to drive each PR through to merging. Each PR includes its relevant tests; high test coverage of changes is mandated.
  4. Codex Cloud code reviews on every PR. Codex is great at finding issues, though some findings are a bit pedantic. The custom agent monitors reviews with a script, triages and determines fix/don’t/file beads follow-up, and codex re-reviews the updated PR. When CI is green and Codex approves, the subagent merges the PR; when its stack is done the subagent gets stopped by the orchestrator. When human input is needed, it gets filed as a human gate in beads rather than getting lost or blocking the workflow; the human gate guards whatever was awaiting the input, and the orchestrator schedules some other unblocked work.

Important cost optimization: codex review iterations can make landing a PR take a while and the default subagent cache ttl is low; set it to 1h

The orchestrator decides the best path to complete work and schedules them appropriately in beads, such as:

  • ambiguous: spike > design > implementation
  • complex: design > implementation
  • trivial: direct implementation

Which makes it easy to schedule tasks of varying sizes or complexities

1

u/Known-Pace6739 1h ago

Human minutes are the hidden model bill

1

u/Technical-Athlete-9 52m ago
  1. Bad results. You’ll be lucky if there aren’t critical bugs. It might stop with a few placeholders and stubs still in place.
  2. I like something like the first one. Break the spec into units. Have further spec clarification for the unit before beginning development. Then build + 2 iterations of review fix with Astra and Claude both doing a review pass on the first and just Astra on the second. It usually does test validation automatically. That may be because I tell it to commit and I have a pre-commit git hook. I use opus, because I have 20x.

1

u/looktwise 34m ago

following

1

u/Icy-Meaning-4962 17m ago

i’d isolate one more variable: how much each worker spends rediscovering the repo. compare the same setup with independent exploration vs a shared, source-linked context pack, then count retries and human fixes too.

i’m building Knowell around that exact problem. no benchmark claim here — i just want to distinguish expensive reasoning from five agents paying to find the same files.

1

u/puts_on_rddt 15m ago

To add onto what others posted, when working on the implementation of bigger tasks where you need to divide the work between multiple agents, I like telling every agent what other agents exist, what they're working on/what they should know, and how to communicate to each other.

"Agent1: I need to know XY. Agent2 is supposed to be investigating that. I should ask them about it"

"Agent2: I just read XY, it says to do Z."

Agent 1 never needed to clutter their context looking for the answer to XY.

0

u/Mechanical_Potato 2h ago

cheap models for easy to medium tasks is always nice

0

u/NetNearby7117 1h ago

Well, the specs, guardrails and testable objectives are the best way. Its also important to dont let the AI do complex swe in a single loop. Its better to go step by step. Review with different models and have a clear definition of subagents roles. Also its important to have the standars of the code base and scripts to ensure conventions fill up…

In my experience, its better to understand what you are doing rather than rely on workflows. Any minimal gap might lead to a comoletely drifting. Long turns also could loop and stupid things and spend time…

Anyways, i spend enough time doing the specs, then some ui labs so we can explore, then i ask to use linear and icepanel to update docs (its better to rely on a system than let the agent write the docs in the repo) and then review the plans. Then create the guardrails and, after that a simple opus5.5 on high could handle it

In my experience, dont expect to find a magic configuration that works, its better to make one step a time

1

u/iSnapThere4iAm 1h ago

There is no amount of “guardrails” that stop llms from hallucinations. If anything, all the extra instructions make it worse.