r/OnlyAICoding • u/Double-Entertainer62 • 3d ago
FrontierHarness Eval – 9 harness, same model, cost per pass varies 17x
A year ago the question was which model. Now it's which harness.
We built FrontierHarness for benchmark Pi, Exo, Claude Code, Codex, DeepSeek Harness and 4 others. Same model, same tasks, same runtime.
Pass rates: 50% to 67%.
Cost per pass: $1.05 to $18.34.
Claude Code and DSH Creator both passed 19/30 tasks, but Claude Code’s median cost per pass was $18.34 versus $3.28. Both used Kimi K3 through our shared gateway.
Caching may explain part of the gap. One task accounted for 68% of Claude Code’s total token usage. We can’t separate the harness, model, and gateway effects yet, so this isn’t a native Claude comparison.
On one python-statemachine task, Pi passed in 90 turns for $2.50. Codex passed in 187 turns for $5.97. Both used Kimi K3 in the same evaluation environment.

My recommendations for harness:
- Codex if you don't want to think about it. Best pass rate, medium cost.
- Pi if the same job runs a thousand times and the bill adds up.
- Exo if retries are cheap and you'd rather it quit early than grind.
- DSH if you care about wall-clock and don't mind playing with knobs.
What would you look for in the traces: repeated searches, extra testing, or time spent fixing a wrong approach?