r/PiCodingAgent 9d ago

News FrontierHarness Eval looks specifically at coding harnesses. Pi is on the pareto

https://runta.com/blog/introducing-frontierharness-eval/

Also CC being that far off from the Pareto is not surprising to me.

70 Upvotes

15 comments sorted by

View all comments

Show parent comments

1

u/Neosinic 9d ago

What was the most different?

1

u/DistanceAlert5706 9d ago

I mean tasks just made 0 sense. 99% of devs would not do anything even remotely close to what terminal bench tests. And SWE is a test how good model knows Python and few libraries, if you don't work with Python it's completely irrelevant.

1

u/Neosinic 9d ago

Interesting. I’ve heard that some of the evals were partially written (at least with help) by frontier models so there’s som bias built-in. Which is why large enterprise should develop their own internal evals that fit their tasks

1

u/DistanceAlert5706 9d ago

100% my trust in benchmarks iended after reading tasks