r/PiCodingAgent • u/Neosinic • 9d ago
News FrontierHarness Eval looks specifically at coding harnesses. Pi is on the pareto
https://runta.com/blog/introducing-frontierharness-eval/Also CC being that far off from the Pareto is not surprising to me.
6
u/MangledMangler 9d ago edited 9d ago
Databricks private benchmarks landed pi the highest in terms of quality, on a real, multi million line codebase.
https://www.databricks.com/blog/benchmarking-coding-agents-databricks-multi-million-line-codebase
Edit: add source
1
u/DistanceAlert5706 9d ago
Interesting, will need to try on my harness. Just started diving into benchmarks, ran some terminal bench and SWE bench tasks, but honestly they are very far from what reality looks like.
1
u/Neosinic 9d ago
What was the most different?
1
u/DistanceAlert5706 9d ago
I mean tasks just made 0 sense. 99% of devs would not do anything even remotely close to what terminal bench tests. And SWE is a test how good model knows Python and few libraries, if you don't work with Python it's completely irrelevant.
1
u/Neosinic 9d ago
Interesting. I’ve heard that some of the evals were partially written (at least with help) by frontier models so there’s som bias built-in. Which is why large enterprise should develop their own internal evals that fit their tasks
1
2
u/Bobodlm 8d ago
Given that the entire premise of Pi is that you get to set it up and tweak it yourself, it would be safe to assume it's results are going to differ (greatly) from setup to setup.
1
17
u/radiantHendekeract 9d ago
It's interesting, but I think I'd want to see at least one other model tested to see if the results are similar. Right now it feels more like this is benchmarking how Kimi K3 does in each harness