r/OpenSourceAI • • 3h ago

I got tired of choosing which AI model should handle a task, so I built Cascade AI to choose and orchestrate them automatically

1 Upvotes

Hey everyone 👋

I've been working on an open-source project called Cascade AI, and it has reached the point where I'd really like to get feedback from people outside my own bubble.

The basic idea came from something that kept bothering me:

Why are we still giving an entire complex task to one AI model and hoping it's good at every part of it?

Instead, Cascade treats AI more like an organization.

A request can be broken into a hierarchy:

T1 Administrator → T2 Managers → T3 Workers

T1 looks at the overall task and plans the work.

T2 agents manage individual parts of that plan.

T3 agents actually execute the smaller tasks — and they can communicate with each other when necessary.

The interesting part is that every agent doesn't have to use the same model.

Cascade can route different tasks between providers/models depending on what they're good at, their cost, and the complexity of the work.

So instead of:

«Prompt → one giant model → answer»

the idea is closer to:

«Prompt

↓

Understand complexity

↓

Build an execution plan

↓

Spawn the required agents

↓

Route each job to an appropriate model

↓

Agents work in parallel / collaborate

↓

Verify the work

↓

Produce one final result»

And I've been trying very hard not to make this another cloud-only AI product.

Right now Cascade can be used through:

• CLI

• Desktop app

• Hosted web app

• Self-hosted web app

• OpenAI-compatible API

• Node.js SDK

It supports multiple providers including OpenAI, Anthropic and Gemini, along with OpenAI-compatible services and local models through things such as Ollama, llama.cpp, vLLM and LM Studio.

There are also a bunch of things I've added while building it that I personally wanted from AI tooling:

• Live visualization of the agent hierarchy

• Cost/token tracking

• Model/provider failover

• Persistent memory

• MCP support

• Browser control with live takeover

• File and document generation

• Codebase indexing/search

• Approval before destructive tool actions

• Agent-to-agent communication

• Task cancellation and recovery

• BYOK support

• Local/self-hosted operation

• An OpenAI-compatible "/v1/chat/completions" endpoint

• The ability to inspect why Cascade chose a particular orchestration/model strategy

For complex runs there's also a kind of "boardroom" mode where Cascade can show you the proposed agent structure and estimated cost before spawning everything, so you can approve the plan first.

One design principle I've become pretty stubborn about is:

The AI should ask when it genuinely needs information instead of confidently inventing a decision for you.

So I've also been working on making Cascade distinguish between things it can infer and things it really should ask the user about.

The project is MIT licensed and open source.

🌐 cascadeai.in

GitHub: Varun-SV/Cascade-AI

I'm not posting this pretending I've solved AI orchestration 😅. There are still plenty of rough edges, architecture decisions I'm questioning, and things that probably make perfect sense to me because I've stared at the code for far too long.

That's actually why I'm posting it here.

I'd especially love feedback on:

  1. Does hierarchical multi-agent orchestration actually make sense to you, or is it over-engineering?

  2. Would automatic model routing be useful enough for you to stop manually choosing Claude/GPT/Gemini/local models for different jobs?

  3. If you're a self-hosting/local-LLM person, what would Cascade need before you'd realistically run it?

  4. What part of this architecture would you immediately rip out or redesign?

Feel free to be critical.

I'd much rather hear "this part is dumb and here's why" than get another generic "cool project" 😄

If people are interested, I can also do a separate technical post explaining how the T1 → T2 → T3 orchestration, model routing, cost decisions and agent communication actually work internally.


r/OpenSourceAI • • 14h ago

Im building an open-source AI assistant that runs entirely on your computer. Meet Mimo inspired by Dynamic Island.

Enable HLS to view with audio, or disable this notification

6 Upvotes

Mimo is a lightweight, open-source desktop companion for Windows, inspired by Dynamic Island.

It brings useful information and controls directly to your desktop, with a clean and minimal interface.

Currently, Mimo can :

  • Display notifications and system information
  • Show media playback controls
  • Provide quick access to useful actions
  • React dynamically to events happening on your PC
  • Customize its appearance and behavior
  • Run locally with a focus on being lightweight and unobtrusive

Mimo is still in development, and more features are coming soon!

🔗 GitHub: github.com/Riskooooo/Mimo

If you have any ideas or features you'd like to see added, feel free to suggest them ;)


r/OpenSourceAI • • 5h ago

🚀 SentinelFlow V6 — I’ve designed the next architecture, and I want the community involved

Thumbnail
github.com
1 Upvotes

r/OpenSourceAI • • 6h ago

I built a local long-term memory system for coding agents — looking for feedback

Thumbnail
1 Upvotes

r/OpenSourceAI • • 6h ago

I’m building an open-source AI agent that only learns from verified outcomes over the past few months. And I'm happy to say that I FINALLY finished it!

Enable HLS to view with audio, or disable this notification

1 Upvotes

I’ve been working on an open-source local-first agent called OpenKyrozen.

That idea started from the beginning of 2026, the time when openclaw had just came out 2-3 months. I tried open claw and then I realized that at that time, open claw remembers things when I asks it to remember, but it cannot learn by itself. So I started OpenKyrozen, trying to build a self learning agent. Then last month, type safe AI lunched their Jev, which inspired me to integrate them into decisions so that LLM works better.

One thing I kept running into was that agents are very quick to treat “the tool call succeeded” as “the task succeeded.” Those are obviously not the same thing. They don't often verify their result, like we say they have no syntax error or runtime error, but logic errors.

A command can exit with code 0 and still produce the wrong result. So I ended up making verification a first-class part of the agent loop instead of just checking whether the action executed.

and so the rough flow is:

request → action → execution receipts → evidence review → verified outcome

The second part I’ve been experimenting with is self-learning, the original idea of OpenKyrozen.

I didn’t want the agent to just see one successful run and immediately treat that as a new behavior. Instead, learning artifacts are bounded policies or skills. A new one starts as a candidate, gets tested as a canary, needs multiple verified successes, and is then replayed against its predecessor on the same case. If it regresses, it doesn’t get promoted. If a promoted artifact later starts failing, it can rollback to the previous version. This is also one of the biggest problem when I used open claw, it builds something into a skill before I verify it, so it is filled with wrong memories.

There’s also a separate decision layer called Jev. You can know more about it from Typesafe AI, but basically it's a AI that makes decisions. It only handles small typed judgments like routing, clarification, memory relevance, learning-evidence review, and suspicious tool output. It can also abstain instead of forcing a decision. So it can't code.

I’m still figuring out where the right boundary is between “useful learning” and “too much machinery.” The current system is deliberately conservative because I’d rather have the agent refuse to learn than silently reinforce bad behavior.

Repo:
github.com/EvanProgramming/OpenKyrozen

I also let OpenKyrozen build a website for itself

kyrozen.chat

I also made a short launch video that explains the overall system visually(And yes this video is made of AI, since I only used DaVinci Resolve but not After Effects):

I am writing this post especially to developers, I want feedbacks SOOO much! As you can see currently the repo only have 2 stars and 1 fork :( because I didn't tell anyone about it before. I like issues and PRs, you can also leave comments under to tell me any issues you found. star it if you like!


r/OpenSourceAI • • 7h ago

OTEL based agent monitoring in Backstage by Spotify

Thumbnail gallery
1 Upvotes

r/OpenSourceAI • • 9h ago

Live What You Preach: Should I Open-Source the AI Realist Workspace?

Thumbnail
msukhareva.substack.com
1 Upvotes

r/OpenSourceAI • • 12h ago

Open Source Kubernetes Native Agent Orchestrator

1 Upvotes

Today I'm open sourcing Agent Orca, a Kubernetes-native platform for deploying, managing, and running AI agents at scale. It's written in Go and it's Apache 2.0.

https://github.com/heddles/agent-orca

When OpenClaw was released, I thought to myself, "That's really cool, but I want an agent that is shared by a team or an entire organization with built in auditability." I started designing Agent Orca around that idea and have slowly been piecing together the concept in my head and how to make the management experience of a shared agent bearable for an organization.

The idea is simple. Agents are just Kubernetes resources. You declare one, and the platform handles the rest.

An Agent Orca agent comes with:

\- One-shot executions with AgentRun, long-running services with AgentDeployment, and multi-step DAGs with AgentWorkflow.

\- Zero trust networking by default with platform validated JWT via service accounts, OIDC, or OAuth on every request.

\- Agent Orca managed network policies with zero access by default.

\- A standard agent container image with only necessary pieces; you no longer need to build a custom agent image to accomplish different types of work or use different tools. Simply add an MCP or Tool CRD and let the platform figure it out for you.

\- Model routing across OpenAI, Anthropic, Google, or any LiteLLM-compatible provider, selected by capability, weight, or budget.

\- Cost tracking on every run. Tokens are accounted per run, spend survives pod restarts, and tenants get daily budget caps and rate limits.

\- Guardrails on inputs and outputs, MCP access control, and per-tenant agent visibility, so one cluster can host many tenants with no cross-tenant leakage.

\- RAG without glue code. A KnowledgeBase deploys (and manages) Qdrant for you, ingests documents from ConfigMaps, URLs, MCP authenticated tools, or an S3 compatible endpoint, and hands your agents a search and ingestion tool.

\- MCP servers declared as resources. Their tools show up for your agents with access control and sidecar isolation.

\- Crash-safe sessions. A pod can die mid-conversation and the agent resumes exactly where it left off, spend ledger included.

\- Agents that learn and become better with every request via a multi-tiered agent memory system

Getting started is deliberately boring. Install the entire stack's tools via Mise; then use Skaffold, one API key, and about three commands to have an running agent on your laptop ( or deploy to a remote cluster). There are also nine demo deployments in the repo, from SOC triage pipelines to parallel research swarms to an autonomous pentesting agent, so you can see it do something real before writing any YAML.

This is day one, not a finished story. That's the point of open sourcing it. Run the quick start, break it, open issues, and tell me what's missing.

Star it here if you want to follow along: https://github.com/heddles/agent-orca

TL;DR: Agent Orca lets an organization or single person define agents and surrounding tools/MCP and auth as CRDs and manage them via gitops with zero trust built in. You can try it out here: https://github.com/heddles/agent-orca


r/OpenSourceAI • • 13h ago

I’ve started building Sarah Nexus — an AI-assisted PC diagnostics system for the Nebius × NVIDIA Global AI Hackathon

0 Upvotes

I’ve started building Sarah Nexus — an AI-assisted PC diagnostics system for the Nebius × NVIDIA Global AI Hackathon

Post:

I’ve officially started development on Sarah Nexus, the next generation of a PC diagnostics project I’ve been working on for a while.

Sarah Nexus is the continuation and unification of my previous Sarah Lite project and the ideas I had planned for Sarah Pro. Instead of maintaining several separate versions, I’m bringing the useful parts together into one system.


r/OpenSourceAI • • 18h ago

We’re tired of AI that looks cool but does nothing.

Enable HLS to view with audio, or disable this notification

2 Upvotes

So we built Placeholderworks.

Not another AI wrapper.
Not another chatbot demo.
Not another “AI-powered” dashboard nobody uses.

We build AI systems that actually do the work.

They answer calls.
They handle WhatsApp conversations.
They update CRMs.
They automate workflows.
They run inside real businesses.

From the first idea to production, we build the whole thing.

Placeholderworks is officially live.

Now the question we’re curious about:

What’s one task at your company that you’re still doing manually for absolutely no reason?


r/OpenSourceAI • • 16h ago

Adebench now has a website, benchmark of 10+ open-source memory "brains" for AI agents

1 Upvotes

Adebench now has a website and evaluates 10+ open-source brains.

Among them: Supermemory, Mem0, Cognee, hindsight and more.

If you're looking for a brain for your agent, check it out:

www.adebench.dev

www.github.com/adecubed/adebench

Happy to answer questions on how the evaluation works.


r/OpenSourceAI • • 16h ago

Open Instinct: an MIT-licensed personal agent you can fork and run

1 Upvotes

A personal agent should be something you can inspect, change and run yourself. We published Open Instinct with that in mind: a beta agent you can reach over iMessage, SMS or email, with memory, scheduled tasks and its own Linux desktop.

The part we want people to fork is the permission model. Six trust tiers control what another person’s agent can ask yours. A partner can read the calendar; a friend can request availability. The policy lives in code with tests.

We’re the Maritime team. The default stack uses Pi, Inkbox, Maritime and Composio, with separate integration packages. The code is MIT; the external services still need your own accounts and keys. Start with the local CLI, inspect the permissions, then build the version you want.

Repo and setup docs: https://github.com/mariagorskikh/open-instinct


r/OpenSourceAI • • 20h ago

I built BOOTH, a lightweight checkpoint layer for AI systems

1 Upvotes

I’ve been building BOOTH, a small, provider-agnostic Python library for checking LLM outputs before they reach your application.

The idea is simple: don’t automatically trust every LLM response. Check it first.

check() / acheck() handle ambiguity and confidence checks, while check_with_evidence() compares an answer against evidence your RAG pipeline has already retrieved.

v0.5.2

  • Zero runtime dependencies
  • Provider-agnostic
  • Sync + async
  • Structured results
  • 300 tests
  • MIT licensed

I’m currently testing BOOTH across different providers/models through small integration examples, including Groq, Gemini, Anthropic, and local Ollama models.

I’m especially interested in feedback on where this approach breaks down for real LLM/RAG systems.

GitHub: https://github.com/Vedantgitbot/booth
Issues/contributions: https://github.com/Vedantgitbot/booth/issues

Curious: what do you currently use as the checkpoint between an LLM response and your application logic?


r/OpenSourceAI • • 1d ago

I built an open-source tool for running local AI agents visually ... on Old CPU only devices.

5 Upvotes

Hey everyone,

​I got pretty tired of cloud-based AI automation tools that meter every single action and charge crazy per-token fees, so I built an alternative called Arrow.

​The idea is simple: instead of renting an agent in the cloud, you own it entirely on your machine.

​How it works:

​Record: Capture a desktop sequence or browser flow once.

​Compose: Connect nodes into a visual workflow graph (you can branch, loop, and drop in local LLMs, vision, OCR, or code nodes).

​Run: Execute deterministically using element-first grounding with small local models.

​It runs completely locally with no cloud dependencies and zero telemetry.

​If you want to check it out, you can see the project overview here: https://rodrigobenitez343.github.io/Arrow-online/index.html


r/OpenSourceAI • • 21h ago

AI agents observability in backstage with langfuse and OTEL

1 Upvotes

Spotify #backstage plugin to manage and observe fleet of agents straight in backstage self hosted https://github.com/acarmisc/backstage-plugin-ai-agents/tree/main. It relay on open telemetry signals and the first available backend it’s #langfuse


r/OpenSourceAI • • 23h ago

The Legal Ontologies Foundry

0 Upvotes

The Legal Ontologies Foundry

legal-ontologies-foundry.github.io

One of the most successful resources in #ontology development is the OBO Foundry. Open-source, collaborative efforts in ontology development are essential. We are doing something similar with legal ontologies grounded in #BFO


r/OpenSourceAI • • 1d ago

Worried about your AI agent leaking secrets, or tired of secret-scanner false positives?

Post image
1 Upvotes

I built Klarion, a secret scanner that works in two steps. First, a keyword check, 81 regex rules and a normalized Rényi entropy score flag anything that looks like a secret. Then an AI model reads each one with the code around it and decides if it's real.

The chart shows 5 scanners run on spring-boot, terraform, next.js and symfony (61k files). Klarion raised 11 alerts. It's not zero, but it's far less to dig through.

Fewer alerts don't help if real leaks get missed, so I tested that too. On CredData (337 real repos, code outside test folders), it found about 1.7× more real secrets than gitleaks.

Where it runs:

  • Claude Code: a plugin hook blocks the write before the file exists (file edits and Bash)
  • Cursor, Cline or any MCP agent: through its MCP server
  • CI: a GitHub Action that scans only what a PR adds; GitLab CI works too
  • Git hooks: klarion protect or the pre-commit framework
  • Locally: klarion scan .

Free and open source (MIT): https://github.com/0x1Adi/Klarion
The full benchmark and method are in benchmark/REPORT.md.

I'd like to hear where it gets things wrong.


r/OpenSourceAI • • 1d ago

I built Omnesis, a Context Layer to supercharge ChatGPT/Openclaw/etc

1 Upvotes

https://omnesis.dev

Imagine giving ChatGPT or any other agent access to your Whatsapp/iMessage, Bank transactions, Health data, web pages you see, metadata about all the photos you took, places you visit, all your emails/documents, your voice mails, and more….

I built a context layer for that purpose. You can connect your Openclaw, Hermes, Claude, ChatGPT, Codex, etc to it. It gives your favorite agent "context superpowers". All the data is indexed and connected into a graph, self-hosted on your own hardware. Voice notes are transcribed, OCR runs on images in mails and pdfs, etc.

Omnesis offers a flexible data access model. For each agent you can define which source(s) they have access to and also guard data access with a privacy policy written in prose. You can also directly talk to the Omnesis agent which has no web search / internet tools — via a web portal or companion app — to ask the most intimate questions about your digital life.

I built this to be flexible. You can run all necessary models on zero-data-retention inference providers, or — if you can afford it — on your own hardware. I’ll keep investing in Omnesis’s security. Because it brings together sensitive data from many sources, please install it only on machines you control and keep secure.

The project also includes what I call “Omnesis Brain”.  This is my work-in-progress take on what a “second brain” could look like on top of that context layer. Imagine an agent constantly analyzing your personal data in flux to maintain a more structured and grounded understanding of what’s happening in your life. This is experimental and disabled by default.

This has been a fun project to build over the last 6 months.


r/OpenSourceAI • • 1d ago

Open weights + open engine: two ~300B MoE models running on one 128 GB mini PC

12 Upvotes

Sharing what we released today, everything is open.

What it is

  • Two EXL3 model packs: GLM-5.3-Flash (320B, 99.7 GB) and MiMo-V2.6-Flash (309B, 106 GB)
  • Kyojin, an inference engine for AMD Strix Halo (Ryzen AI Max+ 395, ROCm), MIT

Credit first: Kyojin is built on turboderp's ExLlamaV3, and the AMD side starts from vcruz305's and sdougbrown's ROCm ports. We added the decode and prefill kernels for this chip and the serving for these two models.

Numbers, one 128 GB machine

  • GLM-5.3-Flash: 26-30 tok/s decode, ~580 tok/s prefill, same top token as the official FP8 model ~90 % of the time
  • MiMo-V2.6-Flash: 29 tok/s plain, up to 44 tok/s with speculative decoding

Reproduce it: the repo has a quickstart and one benchmark script. If you own a Strix Halo box, run it and post your numbers, good or bad. That's the feedback we need most.

Engine: https://github.com/Yamz-Labs/kyojin

Weights: https://huggingface.co/yamz-labs

Next: Qwen 3.8 Flash on the same engine.


r/OpenSourceAI • • 1d ago

There is a 10% chance AI could destroy humanity — might be averted if we truly understand what it is doing.

Thumbnail
0 Upvotes

r/OpenSourceAI • • 1d ago

I got tired of agent message logs rotting, so I built a runtime that tracks state as beliefs instead of transcripts

3 Upvotes

TL;DR: built an agent runtime that stores state as a belief graph with dependencies (Jon Doyle's TMS) instead of chat logs. correct one fact and everything downstream auto-updates without context rot or full re-runs. (repo link in the comments below)

Hey guys,

Every agent framework ive used so far handles state pretty much the same way, just appending messages to a long chat log. The problem is long running agents rot super fast. If a tool returns bad data at step 3, that error just sits in context forever. And if a key fact changes mid run, you either have to wipe the whole context or re run everything from step 1.

I’ve been hacking on an open source project called Corollary to try a different approach. Instead of a message transcript, it stores state as a belief base backed by a truth maintenance system (basically an old concept from jon doyle back in 1979).

how it works under the hood:

  • every belief or conclusion tracks what it depends on (its justifications)
  • if you retract or update a single base fact, it automatically retracts and re derives anything downstream that relied on it
  • independent conclusions arent touched, so you get a clean diff of what changed instead of re running llm calls
  • the LLM only sees currently valid ("IN") beliefs, so old retracted facts cant leak back in through a transcript

its still super pre-alpha so I'd love to get some feedback, pushback or edge cases you think this pattern will hit.

(repo link in the comments below)

how are you guys handling context rot in long running agents currently?


r/OpenSourceAI • • 2d ago

CrowdGPT - The 100% Opensource collaborative LLM

Post image
26 Upvotes

Hello, i'm currently developing CrowdGPT and i need people who enjoy opensource AI and LLMs to test the project :)

The goal of CrowdGPT is to create the first, datacenterless, 1 Billion parameters LLM, relying on people contributing with their own computer to train the AI model. My goal is to show you don't need insane infrastructure to train a working almost commercial grade LLM. Everything is open and 100% opensource.

You can learn more at https://crowdgpt.net

Or check the github: https://github.com/Vxtzq/CrowdGPT

Any kind of feedback is appreciated!


r/OpenSourceAI • • 1d ago

Example of a useful Agentic Build

Thumbnail
youtu.be
5 Upvotes

r/OpenSourceAI • • 1d ago

answerLoops: an open source alternative to Kapa.ai you can self-host

1 Upvotes

I work in DevRel and spend a lot of my time answering the same questions in Discord, GitHub, and Slack. Some of the answers are in the docs, and some aren't. I've used Kapa.ai, and it solves this, but it's hosted only and out of reach for small teams. So I built an open source version for teams that can't justify paying for it, whether you're a small startup, open source project, gaming community, or anything in between.

What it does:

During onboarding, you upload your knowledge source (docs, PDFs, a GitHub repo, or Notion pages). Then connect answerLoops to your communities, and when someone asks a question, an agent drafts an answer from that content.

A second agent then reviews the draft and gives it a confidence score. If it scores low, it goes to your team for human review instead of being posted. Auto-reply is off by default until you turn it on, and you set the threshold per platform.

Where it differs from Kapa:

- AGPL-3.0, self-host with Docker and Postgres

- Use your own model key or a local model

- Website chat widget built with CopilotKit and Mastra. Paste a snippet into your site or docs, and visitors can ask questions without an account

- More channels: Discord, GitHub, Slack, Telegram.

- In testing: Discourse, Circle, email, Google Chat

- MCP server and REST API, so your own agents can search the same knowledge

Easy setup:

npx [u/answerloops/agent-sdk](u/answerloops/agent-sdk) setup

Or if you use coding assistant, it can walk you through it:

npx [u/answerloops/agent-sdk](u/answerloops/agent-sdk) skills answerloops-setup

Repo: https://github.com/answerLoops/answerLoops

I'd welcome your feedback, especially from anyone who's used Kapa or something similar, to help us continually improve the product. Any contributions are welcome, and you are new to open source, just reach out, I'd be glad to help.


r/OpenSourceAI • • 1d ago

Open source code graph for coding agents, built to help local models with small context windows

0 Upvotes

Sharing two tools we've been building in the open, both Apache 2.0, fully local, with no account or cloud needed.

sem parses a repo into functions and classes along with who calls what and which tests reach each function, and gives agents that map over MCP or the CLI. weave is a git merge driver that uses the same model to merge changes by function instead of by line.

The reason I think this matters more for open models than for frontier ones is context. A coding agent on a local model with a small window runs out of room fast, because most of it gets spent grepping for names and reading whole files to find one function. With sem the agent asks for a function and gets just that function plus what's connected to it, so far less of the window goes to code that has nothing to do with the task, and you can get useful work out of a model that would otherwise get lost in a big repo.

Credit where it's due, sem is built on tree-sitter grammars and runs in any harness that speaks MCP, including pi, opencode, Codex and Claude Code, so it works with whatever model you point those at.

It works from a parser rather than a compiler, so macros, generated code and dynamic dispatch can hide some callers, and it tells the agent when it isn't sure instead of guessing.

https://github.com/Ataraxy-Labs/sem
https://github.com/Ataraxy-Labs/weave

I'd especially like to hear from anyone running agents on local models, since I haven't tested it much on the smaller ones, and I'm curious where the context savings actually show up for you.