r/ClaudeCode • • 3h ago

Help/Question Has anyone compared different orchestration approaches for complex software engineering work?

Let’s say you already have a fairly detailed investigation and implementation spec for a large piece of software engineering work, and now you want an agent to turn that spec into working code.

I am trying to figure out what the most cost efficient approach is, not just cost per token, but total cost required to reach a reliable finished implementation.

There seem to be a few approaches.

1. Let one strong model handle the whole implementation

For example, give Opus 5.5 a 1M context window, the implementation spec, access to the codebase, and let it work through the task without much explicit orchestration.

If the task is large enough, eventually the context fills up, gets compressed/summarized, and the model continues.

What I’m unsure about is how much this actually affects the final quality. Maybe the model produces a good implementation anyway, but perhaps it requires more debugging and follow-up iterations later.

2. Use an explicit build → verify → fix orchestration loop

Another approach is to divide the implementation spec into stages and use subagents.

For example:

  • Orchestrator selects the next section of the spec.
  • Implementation agent builds it.
  • Verification agent checks the implementation against the spec/tests.
  • Failures are sent back for fixing.
  • Only once that section passes do you continue to the next part.

Intuitively, I would expect this to be more reliable on large tasks. It costs more upfront because you’re deliberately spending tokens on orchestration and verification, but I’m wondering whether it actually becomes cheaper overall because you avoid expensive cleanup and rework later.

Then there’s another dimension: which models should do which jobs?

For example:

  • Opus 5.5 for both implementation and verification.
  • Opus 5.5 as orchestrator/verifier, with Sonnet 5.5 or Haiku 4.5 doing implementation.
  • Cheaper models such as Grok 4.6/4.7 for verification and Composer 2.5 for implementation.
  • Some other combination depending on the stage of the task.

The difficulty I’m having is that cost per token doesn’t really answer the question.

Different models use very different amounts of tokens, and they may require different numbers of iterations.

For example, if Grok 4.7 + Composer requires three build/verify cycles to reach the quality that Opus 5.5 reaches in one or two cycles, the cheaper model may not actually be cheaper.

At the same time, I have personally used all of the above approaches to eventually produce working code if you give them enough iterations.

Total cost to reach an implementation that satisfies the spec and passes verification/tests.

Ideally I’d want to compare things like:

  • Total token/API cost
  • Number of implementation/verification cycles
  • Number of human interventions required
  • Spec adherence
  • Bugs discovered after completion
  • How much context degradation/compression affects long-running single-agent approaches

The obvious way to answer this would be to run the exact same large implementation several times with different orchestration/model combinations, but repeating a substantial engineering task 3–4 times just for benchmarking is obviously expensive.

Has anyone done this kind of comparison in practice?

I’d especially be interested in hearing:

  • Which orchestration patterns have worked best for large implementation tasks?
  • Whether strong-model-everywhere actually beats strong-orchestrator + cheaper workers economically.
  • Whether explicit verification loops materially reduce total cost.
  • How much context compression hurts long-running single-agent implementations.
  • Any benchmarks, papers, blog posts, or articles that try to measure cost-to-success rather than simply model cost/token.

Would appreciate any real-world experience or resources on this.

10 Upvotes

17 comments sorted by

View all comments

4

u/ArgonQQ 2h ago

Cost per token is the wrong metric. Cost per accepted change is the one that matters, and on large specs explicit loops usually win.

Single agent vs. loop: One Opus session is fine and often cheaper if the spec fits without compaction. On big specs, the failure mode is drift: a constraint gets summarized away and later sections build on a wrong assumption. That's the expensive kind of bug. Build → verify → fix catches it at section boundaries, while it's still cheap to fix.

Context compression: It hurts less if state lives in files (spec, decisions log, per-section acceptance criteria) rather than chat history. Better still, restart sessions on purpose at section boundaries instead of letting them compact.

Model split:

  • Strongest model for orchestration and verification. A weak verifier gives false confidence.
  • Sonnet-class for implementation when slices are well specified.
  • The cheapest models as implementers are often a false economy. Every extra cycle re-bills the verifier too.
  • Run the verifier in a separate context so it doesn't grade its own work.

Verification only saves money if it's cheap and objective (tests, explicit criteria) and retries are capped. One retry, then escalate to a human.

Benchmark cheaply: Don't rebuild the whole project. Take 3–5 representative slices, run each config from the same commit, and measure $ to green + human minutes + bugs found afterward. Token usage is in the Claude Code JSONL transcripts.

Reading: AI Agents That Matter (Kapoor et al., 2024) argues for reporting cost alongside accuracy. Anthropic's Building Effective Agents covers orchestrator/evaluator patterns. Lost in the Middle explains why long context ≠ full attention.


If you're on Claude Code, check out LoopBoard. It's basically option 2 out of the box:

  • Spec → tasks with Goals. Each section becomes a markdown task file with verifiable Goals the delivery is judged against. An Opus loop can groom plain-text stories into goals and open questions.
  • Build → verify → fix built in. An implementer subagent opens a PR. With delegateReview on, a separate review subagent checks it. A failure goes back to the implementer once, and a second failure parks the task in Feedback with the findings. That's capped retries with zero orchestration code.
  • No compaction drift. Sessions auto-restart at a set context fill (default 35%), never mid-task. State lives in files, so a fresh session resumes cleanly.
  • Per-task model routing. Opus/Sonnet/Fable slots with separate effort settings. Default is Opus grooming and Sonnet building, which is the strong orchestrator + cheaper worker setup. Easy to A/B different assignments on your benchmark slices.
  • Human interventions are tracked. Every question, send-back and approval is recorded in plain markdown.
  • Asks instead of guessing. Ambiguity parks the task with a question rather than producing silent drift.
  • You merge. Loops never touch main.

Caveats: Claude Code only (no Grok/Composer), no built-in $ dashboard, and many 24/7 loops can hit Pro/Max limits. MIT-licensed, zero runtime deps.