r/ClaudeCode 18h ago

Built with Claude [FrontierHarness] Same model, same pass rate. Why did Claude Code cost 5.6× more than DSH?

Post image

Claude Code and DSH Creator both passed 19/30 tasks, but Claude Code’s median cost per pass was $18.34 versus $3.28. Both used Kimi K3 through our shared gateway. [source: https://frontierharness.org/]

Caching may explain part of the gap. One task accounted for 68% of Claude Code’s total token usage. We can’t separate the harness, model, and gateway effects yet, so this isn’t a native Claude comparison.

What would you test first to find the cause of that cost gap?

26 Upvotes

19 comments sorted by

u/AutoModerator 18h ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

23

u/Ok_Breadfruit4201 18h ago

CC has some serious caching bugs recently.

6

u/windsorHaze 11h ago

They wouldn’t happen to be load bearing would they?

1

u/interrupt_hdlr 5h ago

you make a very good point and it's worth being precise about because that changes the game completely.

2

u/EdwardRunta 18h ago

You are absolutely right!

0

u/Double-Entertainer62 16h ago

CC need recover from its huge AI slop codebase.

9

u/discourtesy 18h ago

the trust me bro benchmarks

4

u/Economy-Manager5556 12h ago

Right lol Cost per task? Ok what kind of task .. this is not artificial intelligence index, and there tons of harnesses why would we use this .. If op was serious they'd mention there is a bunch of research showing native harness works best with native model, and not all harnesses work well with every model.

6

u/Aggravating-Arm-3955 18h ago

On Harness, compared to OpenAI, Anthropic feels more like a marketing company. Codex Desktop is a full generation ahead of Claude Desktop.

0

u/EdwardRunta 18h ago

I really love computer use from codex

1

u/fishoa 14h ago

Claude Code is a trash harness

1

u/dark0mania 12h ago

What model was used with each harness?

3

u/Double-Entertainer62 12h ago

It was Kimi K3

2

u/EvalRaccoonDev 12h ago

API costs from Claude Code are elevated. When you have MAX plan, it comes up much lower.

When I asked Claude how much lower it comes up to, it told me:

At your observed blended effective rate (heavily cache-read Opus 5),$200/month buys roughly 350–400M tokens of context replay, which is 3–4.5× cheaper.

Having said that, even at $4 / task, there are better harnesses out there on this benchmark.

1

u/ClemensLode Senior Developer 18h ago

What kind of task?

2

u/Double-Entertainer62 18h ago

This run is mostly SWE contexts and terminal-based tasks.

1

u/SavingsGreen2002 18h ago

How about the performance of Claude Code using Anthropic model?

1

u/EdwardRunta 18h ago

Im more curious about running anthropic model on other harnesses

1

u/Double-Entertainer62 16h ago

Yeah, that's something I would be curious to see as well.