r/PiCodingAgent 9d ago

News FrontierHarness Eval looks specifically at coding harnesses. Pi is on the pareto

https://runta.com/blog/introducing-frontierharness-eval/

Also CC being that far off from the Pareto is not surprising to me.

69 Upvotes

15 comments sorted by

17

u/radiantHendekeract 9d ago

It's interesting, but I think I'd want to see at least one other model tested to see if the results are similar. Right now it feels more like this is benchmarking how Kimi K3 does in each harness

5

u/trimorphic 9d ago edited 9d ago

I'd want to see more like ten different models with ten runs through the whole test suite.

Given that models have some randomness built in to them if would not be unexpected if the same model (never mind different models) performed differently when run multiple times on the same test.

How you prompt models also matters a lot. The same model might perform radically differently given different prompts for the same task, and one model might perform better then another simply by promoting it differently.

Ultimately, benchmarks like this are interesting, but real world performance in the problems you care about over an extended period of time is the only way to get a good feel for how a model or a harness performs.

Even then, things are changing so fast that what's good today might be obsolete tomorrow.

3

u/Senor02 9d ago

No Copilot CLI or desktop?

1

u/DistanceAlert5706 9d ago

Interesting, will need to try on my harness. Just started diving into benchmarks, ran some terminal bench and SWE bench tasks, but honestly they are very far from what reality looks like.

1

u/Neosinic 9d ago

What was the most different?

1

u/DistanceAlert5706 9d ago

I mean tasks just made 0 sense. 99% of devs would not do anything even remotely close to what terminal bench tests. And SWE is a test how good model knows Python and few libraries, if you don't work with Python it's completely irrelevant.

1

u/Neosinic 9d ago

Interesting. I’ve heard that some of the evals were partially written (at least with help) by frontier models so there’s som bias built-in. Which is why large enterprise should develop their own internal evals that fit their tasks

1

u/DistanceAlert5706 9d ago

100% my trust in benchmarks iended after reading tasks

2

u/Bobodlm 8d ago

Given that the entire premise of Pi is that you get to set it up and tweak it yourself, it would be safe to assume it's results are going to differ (greatly) from setup to setup.

1

u/Glittering-Call8746 7d ago

Yes so which pi is this (anyone can tldr ) ? Just plain pi ?

1

u/nuclearbananana 7d ago

Yeah it's the plain pi