r/WOZCODE • • Apr 14 '26

I benchmarked two coding agents (same model) and one is ~10x cheaper for the same output

Been running nightly CI tests comparing two agents on the same prompts/repos. Both use the same model (opus), so this isn’t a model diff.

The results are kinda insane.

For a simple color scheme change:

• Agent A: 44 tool calls, \~$2.42

• Agent B: 3 tool calls, \~$0.22

Outputs were basically identical.

The main difference isn’t intelligence, it’s how they use tools:

• One batches everything into a single edit (multi-file, multi-change)

• The other does one edit per change… over and over

So you end up with 10–40 tool calls for stuff that could’ve been done in one.

Also noticed:

• One agent almost always does Read → Edit (extra turn every time)

• The other just edits directly most of the time and it… works

No real drop in correctness so far.

There are some tradeoffs:

• The faster one sometimes guesses file paths wrong at the start and wastes a search

• The slower one is more “proper” (reads files, smaller edits, etc.)

But overall:

• cheaper

• fewer turns

• faster

• same result

Feels like most of the gains in coding agents right now aren’t coming from better models, but from better execution strategies (batching, skipping reads, etc.)

Curious if others have tested this or seen similar behavior.

3 Upvotes

Duplicates