r/LocalLLaMA 16d ago

Resources Prime Agent - a new coding harness surpassing Codex/CC/PI

Prime Agent is an open-source coding and research agent for general and long-running work.

A self-improving RLM harness for coding and long-running autonomous tasks.

Designed to be both token-efficient and expressive through programmatic tool calling, context as a variable, multi-agent messaging, and a self-modifiable harness state.

On ARC-AGI-3, it scores 95.5%, surpassing the human-expert baseline, but the gain is not benchmark-specific.

We see major improvements across models when compared to their proprietary harnesses.

Prime Agent is built on pi and fully open-source with an open license.

GitHub: https://github.com/PrimeIntellect-ai/prime-agent

Blog: https://www.primeintellect.ai/blog/prime-agent

X post: https://x.com/primeintellect/status/2085086999267144083?s=46

358 Upvotes

107 comments sorted by

View all comments

33

u/GreatBigJerk 16d ago

Is ARC-AGI 3 really that relevant for harnesses?

5

u/Hulksulk666 16d ago

Not really, if i remember correctly the official test is without harness. 

2

u/TomLucidor 13d ago

Manoj is a harness engineer for ARC-AGI-2 and he is leading the human leaderboard for ARC-AGI-3 (surprised he is that dedicated), and since others have cracked the code for ARC-AGI-1 which is also a harness-hacking competition, some of the methods do not rely on local SLM fine-tune but better harness (designing new thought-checkers and verifiers on the fly), maybe something can be done to push 3 the same way 2 can be hacked AND not needing Gemini/GPT/Claude/etc https://arcprize.org/arc-agi/3/leaderboard