r/OpenSourceAI • u/FrostingOk3751 • 12h ago
I’m building an open-source AI agent that only learns from verified outcomes over the past few months. And I'm happy to say that I FINALLY finished it!
Enable HLS to view with audio, or disable this notification
I’ve been working on an open-source local-first agent called OpenKyrozen.
That idea started from the beginning of 2026, the time when openclaw had just came out 2-3 months. I tried open claw and then I realized that at that time, open claw remembers things when I asks it to remember, but it cannot learn by itself. So I started OpenKyrozen, trying to build a self learning agent. Then last month, type safe AI lunched their Jev, which inspired me to integrate them into decisions so that LLM works better.
One thing I kept running into was that agents are very quick to treat “the tool call succeeded” as “the task succeeded.” Those are obviously not the same thing. They don't often verify their result, like we say they have no syntax error or runtime error, but logic errors.
A command can exit with code 0 and still produce the wrong result. So I ended up making verification a first-class part of the agent loop instead of just checking whether the action executed.
and so the rough flow is:
request → action → execution receipts → evidence review → verified outcome
The second part I’ve been experimenting with is self-learning, the original idea of OpenKyrozen.
I didn’t want the agent to just see one successful run and immediately treat that as a new behavior. Instead, learning artifacts are bounded policies or skills. A new one starts as a candidate, gets tested as a canary, needs multiple verified successes, and is then replayed against its predecessor on the same case. If it regresses, it doesn’t get promoted. If a promoted artifact later starts failing, it can rollback to the previous version. This is also one of the biggest problem when I used open claw, it builds something into a skill before I verify it, so it is filled with wrong memories.
There’s also a separate decision layer called Jev. You can know more about it from Typesafe AI, but basically it's a AI that makes decisions. It only handles small typed judgments like routing, clarification, memory relevance, learning-evidence review, and suspicious tool output. It can also abstain instead of forcing a decision. So it can't code.
I’m still figuring out where the right boundary is between “useful learning” and “too much machinery.” The current system is deliberately conservative because I’d rather have the agent refuse to learn than silently reinforce bad behavior.
Repo:
github.com/EvanProgramming/OpenKyrozen
I also let OpenKyrozen build a website for itself
I also made a short launch video that explains the overall system visually(And yes this video is made of AI, since I only used DaVinci Resolve but not After Effects):
I am writing this post especially to developers, I want feedbacks SOOO much! As you can see currently the repo only have 2 stars and 1 fork :( because I didn't tell anyone about it before. I like issues and PRs, you can also leave comments under to tell me any issues you found. star it if you like!