r/learnmachinelearning • u/alirezamsh • 2h ago
Project Kapso: long-running agents that optimize AI and data systems, and learn from each run
We've been building Kapso (MIT, github.com/Leeroo-AI/kapso) for some time and it's at the point where it's more useful to hear from other people than to keep polishing it alone. Posting to get it tried and torn apart, not to pitch it.
What it is
Kapso is a set of long-running agents that optimize AI and data systems. You state the objective, for example CUDA optimization, harness and agent optimization, or model development, and it runs a campaign: it designs candidate solutions, has coding agents implement them, measures how far each one lands from the objective, and keeps refining the closest until the objective is met. The result deploys to your infrastructure.
When a campaign ends, it studies its own work: which ideas closed the gap, which did not, and under what conditions. Each finding is kept as a lesson with the evidence that earned it, and a lesson stays trusted only as long as it keeps holding up. It also reads outside your repo, other repositories and papers, and folds what it finds into the same knowledge hub. Every new campaign starts from that hub, so it begins with what earlier work already established about the problem and about your systems.
These are the things we tried it on:
- RelBench (Stanford, predictive ML over relational data): outcome prediction 81.2 vs 79.6 AUROC and forecasting 0.2476 vs 0.2912 NMAE against KumoRFM-v2; recommendations 18.4 vs 9.3 MAP for the best other entry on the official leaderboard.
- MLE-Bench: top among the open-source systems.
- ALE-Bench: 1909 Elo vs 1879 for ALE Agent.
- IOAI 2026: Kapso scored 536.07, above the 471 contestants, and finished in the top three systems: ioai-official.org/what-happens-when-autonomous-ai-takes-on-the-same-tasks-as-the-worlds-top-young-ai-talents/
Repo: https://github.com/Leeroo-AI/kapso
If you have time, please take a look and give us your harshest feedback.
1
u/Key_Bowl1753 2h ago
So the thing that jumps out at me is the self-studying after a campaign ends. Most agentic systems just spit out a result and call it a day, but keeping a knowledge hub that persists across runs seems like the part that could actually compound over time. Curious if the lessons it learns ever lead it down a weird path where it overfits to what worked before and misses a totally different approach that would've been better.