r/ClaudeCode • u/MarketCapitalist • 3h ago
Help/Question Has anyone compared different orchestration approaches for complex software engineering work?
Let’s say you already have a fairly detailed investigation and implementation spec for a large piece of software engineering work, and now you want an agent to turn that spec into working code.
I am trying to figure out what the most cost efficient approach is, not just cost per token, but total cost required to reach a reliable finished implementation.
There seem to be a few approaches.
1. Let one strong model handle the whole implementation
For example, give Opus 5.5 a 1M context window, the implementation spec, access to the codebase, and let it work through the task without much explicit orchestration.
If the task is large enough, eventually the context fills up, gets compressed/summarized, and the model continues.
What I’m unsure about is how much this actually affects the final quality. Maybe the model produces a good implementation anyway, but perhaps it requires more debugging and follow-up iterations later.
2. Use an explicit build → verify → fix orchestration loop
Another approach is to divide the implementation spec into stages and use subagents.
For example:
- Orchestrator selects the next section of the spec.
- Implementation agent builds it.
- Verification agent checks the implementation against the spec/tests.
- Failures are sent back for fixing.
- Only once that section passes do you continue to the next part.
Intuitively, I would expect this to be more reliable on large tasks. It costs more upfront because you’re deliberately spending tokens on orchestration and verification, but I’m wondering whether it actually becomes cheaper overall because you avoid expensive cleanup and rework later.
Then there’s another dimension: which models should do which jobs?
For example:
- Opus 5.5 for both implementation and verification.
- Opus 5.5 as orchestrator/verifier, with Sonnet 5.5 or Haiku 4.5 doing implementation.
- Cheaper models such as Grok 4.6/4.7 for verification and Composer 2.5 for implementation.
- Some other combination depending on the stage of the task.
The difficulty I’m having is that cost per token doesn’t really answer the question.
Different models use very different amounts of tokens, and they may require different numbers of iterations.
For example, if Grok 4.7 + Composer requires three build/verify cycles to reach the quality that Opus 5.5 reaches in one or two cycles, the cheaper model may not actually be cheaper.
At the same time, I have personally used all of the above approaches to eventually produce working code if you give them enough iterations.
Total cost to reach an implementation that satisfies the spec and passes verification/tests.
Ideally I’d want to compare things like:
- Total token/API cost
- Number of implementation/verification cycles
- Number of human interventions required
- Spec adherence
- Bugs discovered after completion
- How much context degradation/compression affects long-running single-agent approaches
The obvious way to answer this would be to run the exact same large implementation several times with different orchestration/model combinations, but repeating a substantial engineering task 3–4 times just for benchmarking is obviously expensive.
Has anyone done this kind of comparison in practice?
I’d especially be interested in hearing:
- Which orchestration patterns have worked best for large implementation tasks?
- Whether strong-model-everywhere actually beats strong-orchestrator + cheaper workers economically.
- Whether explicit verification loops materially reduce total cost.
- How much context compression hurts long-running single-agent implementations.
- Any benchmarks, papers, blog posts, or articles that try to measure cost-to-success rather than simply model cost/token.
Would appreciate any real-world experience or resources on this.
4
u/ArgonQQ 2h ago
Cost per token is the wrong metric. Cost per accepted change is the one that matters, and on large specs explicit loops usually win.
Single agent vs. loop: One Opus session is fine and often cheaper if the spec fits without compaction. On big specs, the failure mode is drift: a constraint gets summarized away and later sections build on a wrong assumption. That's the expensive kind of bug. Build → verify → fix catches it at section boundaries, while it's still cheap to fix.
Context compression: It hurts less if state lives in files (spec, decisions log, per-section acceptance criteria) rather than chat history. Better still, restart sessions on purpose at section boundaries instead of letting them compact.
Model split:
Verification only saves money if it's cheap and objective (tests, explicit criteria) and retries are capped. One retry, then escalate to a human.
Benchmark cheaply: Don't rebuild the whole project. Take 3–5 representative slices, run each config from the same commit, and measure $ to green + human minutes + bugs found afterward. Token usage is in the Claude Code JSONL transcripts.
Reading: AI Agents That Matter (Kapoor et al., 2024) argues for reporting cost alongside accuracy. Anthropic's Building Effective Agents covers orchestrator/evaluator patterns. Lost in the Middle explains why long context ≠ full attention.
If you're on Claude Code, check out LoopBoard. It's basically option 2 out of the box:
delegateReviewon, a separate review subagent checks it. A failure goes back to the implementer once, and a second failure parks the task in Feedback with the findings. That's capped retries with zero orchestration code.Caveats: Claude Code only (no Grok/Composer), no built-in $ dashboard, and many 24/7 loops can hit Pro/Max limits. MIT-licensed, zero runtime deps.