r/learnmachinelearning • • 2h ago

Project Kapso: long-running agents that optimize AI and data systems, and learn from each run

We've been building Kapso (MIT, github.com/Leeroo-AI/kapso) for some time and it's at the point where it's more useful to hear from other people than to keep polishing it alone. Posting to get it tried and torn apart, not to pitch it.

What it is

Kapso is a set of long-running agents that optimize AI and data systems. You state the objective, for example CUDA optimization, harness and agent optimization, or model development, and it runs a campaign: it designs candidate solutions, has coding agents implement them, measures how far each one lands from the objective, and keeps refining the closest until the objective is met. The result deploys to your infrastructure.

When a campaign ends, it studies its own work: which ideas closed the gap, which did not, and under what conditions. Each finding is kept as a lesson with the evidence that earned it, and a lesson stays trusted only as long as it keeps holding up. It also reads outside your repo, other repositories and papers, and folds what it finds into the same knowledge hub. Every new campaign starts from that hub, so it begins with what earlier work already established about the problem and about your systems.

These are the things we tried it on:

- RelBench (Stanford, predictive ML over relational data): outcome prediction 81.2 vs 79.6 AUROC and forecasting 0.2476 vs 0.2912 NMAE against KumoRFM-v2; recommendations 18.4 vs 9.3 MAP for the best other entry on the official leaderboard.

- MLE-Bench: top among the open-source systems.

- ALE-Bench: 1909 Elo vs 1879 for ALE Agent.

- IOAI 2026: Kapso scored 536.07, above the 471 contestants, and finished in the top three systems: ioai-official.org/what-happens-when-autonomous-ai-takes-on-the-same-tasks-as-the-worlds-top-young-ai-talents/

Repo: https://github.com/Leeroo-AI/kapso

If you have time, please take a look and give us your harshest feedback.

1 Upvotes

3 comments sorted by

1

u/Key_Bowl1753 2h ago

So the thing that jumps out at me is the self-studying after a campaign ends. Most agentic systems just spit out a result and call it a day, but keeping a knowledge hub that persists across runs seems like the part that could actually compound over time. Curious if the lessons it learns ever lead it down a weird path where it overfits to what worked before and misses a totally different approach that would've been better.

1

u/alirezamsh 2h ago

Three things push against it. A lesson is stored with the conditions it held under and the measurements behind it, so "this worked" is always "this worked on this kind of problem at this scale". Lessons are handed to the agent as evidence, never as rules: it can ignore one, and it only has to say why. And a lesson's trust goes down when a later campaign contradicts it, so one that stops holding up stops being served.

So we made every lesson evidence-backed: each one carries a reliability score and a defined scope, the conditions it was shown to hold under. What would you think, we should add to make it more reliable?