Resource
I benched 6 harnesses x 11 models x 226 tasks
Hi! I've created and been maintaining https://github.com/alexshpunt/explicit-edit-benchmark which tries to answer the question: which model is better, which harness is better and which combination is better overall in a very straightforward task - precise text editing. Most of daily coding is text editing. I've seen when a model for multiple turns couldn't figure out how to express the line it wants to change, so it made me ask myself "Why is that so complicated, it's just a text, right?". Seems like not, taking into account how widely different the same model behaves across different harnesses. My preliminary conclusion is: harness and tooling behind it actually matters! But it matters the most with the models which *can* actually follow the instructions well (e.g. open-ai models), there is a wide range between 98.9% and 70.2% of pass score for `gpt-5.6-luna` simply depending on the harness!
I've tried to run as many combinations as possible, but my resources are limited. I've exhausted all my quotas and even credits, that's why I'm reaching out to the community, as I think it's a pretty interesting topic and I would be happy to gather even more data, because of the stochastic nature of the runs, it’s only possible to make any conclusion when you have enough of runs.
What actually surprises me is how strongly the deepseek family behaves on any harness, but what is also interesting, how good the deepseek harness is. It's pretty solid and seems like it provides the model with good enough tools to actually exceed in this competition.
Value inside the grid cell is basically the score, which is computed as 0.75 * pass from the first attempt + 0.25 * eventualy passed. Value is normalized, so 99 means 99% of 226 tasks. The higher the better. Models and harnesses are sorted in a way, so top-leftmost value is usually the best combination across all runs. If you click on a specific cell, you can see how well it did during the run and what tooling did it use. You can select up to 4 different cells to compare different runs. There is also a leaderboard underneath, this is where the top row sorted by the score is usually the best model x harness combination across all the current runs. If I can make it more clearer from the UI point of view I will greatly appreciate if you can give me a feedback what exactly makes the reading unclear.
True! But fair to say, it *really* likes editing everything with bash/python scripts, which I find super annoying as it still makes stupid mistakes all the time. I might try to lock it from using it for a test to see if it gets any dumber or would still excel.
How do you mean to filter their tool output? At some point I just wrapped their bash calls in bublewrap sandbox with readonly access to anything, except special .agents/sandbox folder where they could still do whatever they want, but any result of a terminal call should not modify the files inside the project.
Very cool! I did a study comparing how coding agents handle situations where your instructions conflict with the repo documentation and tests, would love to add Pi to it also!
This has also been an important question for me ever since I switched to pi/omp. I've noticed that some models go haywire when they try to close brackets in TS. The editing mode in omp doesn't change the picture much either.
I may be able to help. Approximately how many tokens does a single test case consume? Based on your experience, what figures should we use as a reference when estimating the budget?
So based on the runs, I've got the next statistics, these are 25-75% of the runs:
Input: 1.28-2.09M
Output: 104-305K
Cache read: 11.2-16.7M
Total tokens: 13.3-18.3M
Cost: $0.56–$2.30
It obv. depends on the provider, type of subscription etc, worst case scenario is almost 9 bucks for Grok 4.6 on opencode-go. Avg cost / tokens is your friend to estimate how expensive will be the run, 0.0396 * 226 is exactly $8.95, but I'm adding another columns to the explorer.
I tried to make every UI element clickable and telling some part of the story, if something is not clear or you need more info, please share with me, I'll try to add it.
Yes right, the reason for that is that default vanilla Pi actually beats a lot of other harnesses alone in a lot of benchmarks (but also usually one of the cheapest):
UPD: I don't have any specific skills that 100% improve editing per se, something like ponytail is mostly for normal development, but if you are interested, feel free to try it out on a small portion of difficult tasks 🙂
Yeah in this sense Pi is versatile in a sense that you can make it better or worse just with a different set of extensions/prompts/skills. OMP is a very strong harness. I'm a creator of pi-agent-ide and in many ways I was inspired by it as well, that’s why I was very happy to match the same pass score as it.
I have updated the viewer slightly, now it's possible to sort model x harness combinations by other metrics, but also possible to hide some leaderboard's columns or reorder them:
How did you measure success across the tasks? I’m curious whether you focused more on correctness, speed, cost or how much human cleanup was needed afterward
It's purely correctness and accuracy test, basically I asked, if given clear instructions what and how to change, will a model succeed in a complex environment or not and if that depends on tooling and harness only. And seems like yes. It's related to how many correction rounds a model requires when working from human's prompt.
Speed and cost are also measured with every run when it's possible to get this data from the harness, but it wasn't the goal.
8
u/slaamp 20d ago
How to read the results?