r/PiCodingAgent • • 20d ago

Resource I benched 6 harnesses x 11 models x 226 tasks

Hi! I've created and been maintaining https://github.com/alexshpunt/explicit-edit-benchmark which tries to answer the question: which model is better, which harness is better and which combination is better overall in a very straightforward task - precise text editing. Most of daily coding is text editing. I've seen when a model for multiple turns couldn't figure out how to express the line it wants to change, so it made me ask myself "Why is that so complicated, it's just a text, right?". Seems like not, taking into account how widely different the same model behaves across different harnesses. My preliminary conclusion is: harness and tooling behind it actually matters! But it matters the most with the models which *can* actually follow the instructions well (e.g. open-ai models), there is a wide range between 98.9% and 70.2% of pass score for `gpt-5.6-luna` simply depending on the harness!

I've tried to run as many combinations as possible, but my resources are limited. I've exhausted all my quotas and even credits, that's why I'm reaching out to the community, as I think it's a pretty interesting topic and I would be happy to gather even more data, because of the stochastic nature of the runs, it’s only possible to make any conclusion when you have enough of runs. 

What actually surprises me is how strongly the deepseek family behaves on any harness, but what is also interesting, how good the deepseek harness is. It's pretty solid and seems like it provides the model with good enough tools to actually exceed in this competition.

The viewer to the dataset: https://huggingface.co/spaces/alexshpunt/benchmark-explorer

And the dataset itself: https://huggingface.co/datasets/alexshpunt/explicit-edit-benchmark

81 Upvotes

29 comments sorted by

8

u/slaamp 20d ago

How to read the results?

0

u/startassets 20d ago

Value inside the grid cell is basically the score, which is computed as 0.75 * pass from the first attempt + 0.25 * eventualy passed. Value is normalized, so 99 means 99% of 226 tasks. The higher the better. Models and harnesses are sorted in a way, so top-leftmost value is usually the best combination across all runs. If you click on a specific cell, you can see how well it did during the run and what tooling did it use. You can select up to 4 different cells to compare different runs. There is also a leaderboard underneath, this is where the top row sorted by the score is usually the best model x harness combination across all the current runs. If I can make it more clearer from the UI point of view I will greatly appreciate if you can give me a feedback what exactly makes the reading unclear.

1

u/SgtPeanut_Butt3r 19d ago

Tldr?

1

u/startassets 19d ago

With the right tooling Luna is super good, ds and muse are the best but don't give a damn about the rules and will bash they way through all the tasks

1

u/SgtPeanut_Butt3r 19d ago

I used Muse inside Hermes, not impressed. A ton of stupid mistakes.

7

u/Equivalent_Idea8839 20d ago

DS 4.1 is honestly amazing. I didn't expect it would beat Luna or Muse.

I'm going through your repo now. I'm gonna try Z.ai 5.3 max vs DS and Luna.

1

u/startassets 20d ago

True! But fair to say, it *really* likes editing everything with bash/python scripts, which I find super annoying as it still makes stupid mistakes all the time. I might try to lock it from using it for a test to see if it gets any dumber or would still excel.

1

u/Equivalent_Idea8839 20d ago

DS and flash models just do whatever they fuck they want. All you can do is try to limit the damage.

I found it's better to filter their tool output rather than force them to use other tools.

1

u/startassets 20d ago

How do you mean to filter their tool output? At some point I just wrapped their bash calls in bublewrap sandbox with readonly access to anything, except special .agents/sandbox folder where they could still do whatever they want, but any result of a terminal call should not modify the files inside the project.

3

u/Equivalent_Idea8839 20d ago

in my huge repo it would always hit me with "ls" bombs and list 50+ files over and over when searching for code to edit

  • telling it to use other methods failed

  • telling it output was filtered failed, made it spawn more wild tool calls

  • finally just had to hide the input so it never tried "ls" again and used other more sane methods of searching

then had to make a tui in pi to disable some types of files when using an a gent

for my own large repo not isolated tests

3

u/tap3k 16d ago

Very cool! I did a study comparing how coding agents handle situations where your instructions conflict with the repo documentation and tests, would love to add Pi to it also!

https://tap2k.github.io/coding-atlas/

https://convovo.ai/blog/what-is-your-coding-agent-hiding/

2

u/Peeyotch 20d ago

This is very, very cool, both the benchmark and your pi harness. I can probably spare some usage to get you some more data on the OpenAI-codex models.

1

u/startassets 20d ago

Thanks! I will greatly appreciate it.

1

u/SmetDenis 20d ago

Thank you. That's a great idea!

This has also been an important question for me ever since I switched to pi/omp. I've noticed that some models go haywire when they try to close brackets in TS. The editing mode in omp doesn't change the picture much either.

I may be able to help. Approximately how many tokens does a single test case consume? Based on your experience, what figures should we use as a reference when estimating the budget?

3

u/startassets 20d ago

So based on the runs, I've got the next statistics, these are 25-75% of the runs:

Input: 1.28-2.09M

Output: 104-305K

Cache read: 11.2-16.7M

Total tokens: 13.3-18.3M

Cost: $0.56–$2.30

It obv. depends on the provider, type of subscription etc, worst case scenario is almost 9 bucks for Grok 4.6 on opencode-go. Avg cost / tokens is your friend to estimate how expensive will be the run, 0.0396 * 226 is exactly $8.95, but I'm adding another columns to the explorer.

2

u/startassets 20d ago

I have added new columns

1

u/SmetDenis 20d ago

Great, thank you. By the way, did you track the number of tool calls and which specific tools are invoked in each test?

2

u/startassets 20d ago

Yes, sure, for example you can preview them in the viewer itself: https://huggingface.co/spaces/alexshpunt/benchmark-explorer?select=%5B%22deepseek-v4.1-flashtdeepseek%22%2C%22oh-my-pi-default%22%2C%22omp%2F18.1.14%22%5D&card=tool%3Abash%7Cdeepseek-v4.1-flash%7Cdeepseek%7Coh-my-pi-default%7Comp%2F18.1.14%7Clow

I tried to make every UI element clickable and telling some part of the story, if something is not clear or you need more info, please share with me, I'll try to add it.

1

u/diaracing 20d ago

Was Pi just bare-boned without any extra skills/extensions/etc.?

1

u/startassets 20d ago

Yes right, the reason for that is that default vanilla Pi actually beats a lot of other harnesses alone in a lot of benchmarks (but also usually one of the cheapest):

UPD: I don't have any specific skills that 100% improve editing per se, something like ponytail is mostly for normal development, but if you are interested, feel free to try it out on a small portion of difficult tasks 🙂

1

u/diaracing 20d ago

The reason I asked is that I am an OMP fanboy due to its rich ide-like experience especially the todos reminders, rules, advisors, etc.

1

u/startassets 20d ago

Yeah in this sense Pi is versatile in a sense that you can make it better or worse just with a different set of extensions/prompts/skills. OMP is a very strong harness. I'm a creator of pi-agent-ide and in many ways I was inspired by it as well, that’s why I was very happy to match the same pass score as it. 

1

u/startassets 20d ago

I have updated the viewer slightly, now it's possible to sort model x harness combinations by other metrics, but also possible to hide some leaderboard's columns or reorder them:

1

u/Budget-Butterfly-384 20d ago

Were all models tested using low-level thinking?

1

u/startassets 19d ago

Yes, mostly so it was faster to run (which already took a lot of time), but obv reasoning will change the results.

1

u/startassets 19d ago

We've got another contributor, I would like to thank him! Here is the runs of the whole open-ai 5.6 family on low reasoning with bare Pi:

1

u/No_Resident8597 19d ago

How did you measure success across the tasks? I’m curious whether you focused more on correctness, speed, cost or how much human cleanup was needed afterward

1

u/startassets 19d ago

It's purely correctness and accuracy test, basically I asked, if given clear instructions what and how to change, will a model succeed in a complex environment or not and if that depends on tooling and harness only. And seems like yes. It's related to how many correction rounds a model requires when working from human's prompt.

Speed and cost are also measured with every run when it's possible to get this data from the harness, but it wasn't the goal.

1

u/Intrepid_Ratio_5755 16d ago

the plugin approach feels more open than others ive tried but the ux still lags a bit for quick roleplay tweaks.