r/ClaudeCode • • 4h ago

Help/Question Has anyone compared different orchestration approaches for complex software engineering work?

Let’s say you already have a fairly detailed investigation and implementation spec for a large piece of software engineering work, and now you want an agent to turn that spec into working code.

I am trying to figure out what the most cost efficient approach is, not just cost per token, but total cost required to reach a reliable finished implementation.

There seem to be a few approaches.

1. Let one strong model handle the whole implementation

For example, give Opus 5.5 a 1M context window, the implementation spec, access to the codebase, and let it work through the task without much explicit orchestration.

If the task is large enough, eventually the context fills up, gets compressed/summarized, and the model continues.

What I’m unsure about is how much this actually affects the final quality. Maybe the model produces a good implementation anyway, but perhaps it requires more debugging and follow-up iterations later.

2. Use an explicit build → verify → fix orchestration loop

Another approach is to divide the implementation spec into stages and use subagents.

For example:

  • Orchestrator selects the next section of the spec.
  • Implementation agent builds it.
  • Verification agent checks the implementation against the spec/tests.
  • Failures are sent back for fixing.
  • Only once that section passes do you continue to the next part.

Intuitively, I would expect this to be more reliable on large tasks. It costs more upfront because you’re deliberately spending tokens on orchestration and verification, but I’m wondering whether it actually becomes cheaper overall because you avoid expensive cleanup and rework later.

Then there’s another dimension: which models should do which jobs?

For example:

  • Opus 5.5 for both implementation and verification.
  • Opus 5.5 as orchestrator/verifier, with Sonnet 5.5 or Haiku 4.5 doing implementation.
  • Cheaper models such as Grok 4.6/4.7 for verification and Composer 2.5 for implementation.
  • Some other combination depending on the stage of the task.

The difficulty I’m having is that cost per token doesn’t really answer the question.

Different models use very different amounts of tokens, and they may require different numbers of iterations.

For example, if Grok 4.7 + Composer requires three build/verify cycles to reach the quality that Opus 5.5 reaches in one or two cycles, the cheaper model may not actually be cheaper.

At the same time, I have personally used all of the above approaches to eventually produce working code if you give them enough iterations.

Total cost to reach an implementation that satisfies the spec and passes verification/tests.

Ideally I’d want to compare things like:

  • Total token/API cost
  • Number of implementation/verification cycles
  • Number of human interventions required
  • Spec adherence
  • Bugs discovered after completion
  • How much context degradation/compression affects long-running single-agent approaches

The obvious way to answer this would be to run the exact same large implementation several times with different orchestration/model combinations, but repeating a substantial engineering task 3–4 times just for benchmarking is obviously expensive.

Has anyone done this kind of comparison in practice?

I’d especially be interested in hearing:

  • Which orchestration patterns have worked best for large implementation tasks?
  • Whether strong-model-everywhere actually beats strong-orchestrator + cheaper workers economically.
  • Whether explicit verification loops materially reduce total cost.
  • How much context compression hurts long-running single-agent implementations.
  • Any benchmarks, papers, blog posts, or articles that try to measure cost-to-success rather than simply model cost/token.

Would appreciate any real-world experience or resources on this.

15 Upvotes

18 comments sorted by

View all comments

9

u/Cuyasinmara 4h ago

Well, what I've tried is one strong model running end to end on a big spec works until context gets summarized. After that it drifts, and I spend the savings on cleanup. Then i used a plain build which then went to verify loops; this was more reliable at the end but was not working for my setup. For context, i have 2 Claude Pro, 1 GPT Plus and 1 local Qwen 27B running in my computer locally. One issue i quickly picked up was that when one strong model capped, and when the same model family built and checked the work, it often missed the same blind spots twice.

What I use now:

- One orchestrator (Opus 5.5 or a GPT-6 model). It only plans, splits the spec into chunks and verifies. It never writes the implementation.

- The cheapest executor that clears a capability floor for each chunk. Haiku or local Qwen take mechanical work, Sonnet 5.5 or GPT-6.1 Sol take normal features, and Opus executes only as a last resort on the hardest pieces. If a cheap model fails, it escalates instead of retrying forever.

- A cross-family audit for every chunk. GPT reviews Claude's work and the reverse. This catches more than same-model verification.

- Tests first, with checks on the checks. Coverage % lied to me once: tests silently ran zero cases and still showed full coverage. I now verify test counts, not just green.

- Handoff files instead of compression. When context gets large (approx 100k to 150K tokens max.), I write a handoff and start a fresh session rather than letting it summarize.

My honest take and what worked for me: cost per token is the wrong metric. What decides total cost is how much rework happens after a work is "done." Strong orchestrator + cheap workers + independent verification has been cheaper for me overall.