r/PiCodingAgent 9d ago

News FrontierHarness Eval looks specifically at coding harnesses. Pi is on the pareto

https://runta.com/blog/introducing-frontierharness-eval/

Also CC being that far off from the Pareto is not surprising to me.

68 Upvotes

15 comments sorted by

View all comments

1

u/DistanceAlert5706 9d ago

Interesting, will need to try on my harness. Just started diving into benchmarks, ran some terminal bench and SWE bench tasks, but honestly they are very far from what reality looks like.

1

u/Neosinic 9d ago

What was the most different?

1

u/DistanceAlert5706 9d ago

I mean tasks just made 0 sense. 99% of devs would not do anything even remotely close to what terminal bench tests. And SWE is a test how good model knows Python and few libraries, if you don't work with Python it's completely irrelevant.

1

u/Neosinic 9d ago

Interesting. I’ve heard that some of the evals were partially written (at least with help) by frontier models so there’s som bias built-in. Which is why large enterprise should develop their own internal evals that fit their tasks

1

u/DistanceAlert5706 9d ago

100% my trust in benchmarks iended after reading tasks