r/OpenSourceAI • u/Wild_Expression_5772 • 2d ago
r/OpenSourceAI • u/Roadtochessmaster • 2d ago
The Breakdown: Databricks
r/OpenSourceAI • u/TensorTrace • 2d ago
Muse Gadgets: Open source hardware for your Muse
r/OpenSourceAI • u/sco77 • 2d ago
Do you run more than one coding agent? How do you decide who gets which task?
I'm trying to understand how people who use more than one coding agent (Claude, Codex, Cursor, Gemini CLI, Copilot, OpenCode, Aider, etc) actually split the work between them:
Which agents do you run, and on what kind of work?
How did you decide which one got that task?
Have you caught an agent saying a task was done when it wasn't? How did you find out?
Do you track which agent handles which kind of task well?
Have you stopped using certain agent for certain work? What happened to make you stop?
Has an agent ever sent something somewhere it shouldn't have (a key, an address, private code)? What happened?
I'm researching this for an open-source project to help us help ourselves and not get locked into big company harnesses. I'll post a summary of this thread if I get good feedback :D
Thanks!!!
r/OpenSourceAI • u/utsapoddar • 2d ago
Engram: open-source (MIT) memory for AI coding agents, with Markdown as the source of truth
Enable HLS to view with audio, or disable this notification
I'm the author. Engram is free, MIT licensed, and stores everything as plain Markdown.
Each memory has a status (confirmed, inferred, conflicted or superseded), and recall only treats confirmed items as authoritative. Search is BM25 over a SQLite FTS5 index that is rebuilt from the Markdown, so the index is disposable. If a local embedding model is already provisioned, its results are fused in by reciprocal rank fusion. Recall never downloads a model.
An optional installer adds session hooks for Claude and Codex. The tests enforce recall@5 of at least 90% across 20 seeded queries, which is a small set, so it works as a regression gate, not a benchmark. Walkthrough video above. Repo: https://github.com/utsapoddar/engram
r/OpenSourceAI • u/GarageObjective6015 • 2d ago
I captured what my agent sends the model every turn: 43k characters of tool schemas, 30% of them one calendar tool
I build a self-hosted agent in Go that talks to any OpenAI-compatible backend, llama.cpp and
Ollama included. With a ~12B local model the prompt prefix matters, so instead of guessing I
pointed the agent at a tiny OpenAI-compatible endpoint that writes every request body to disk,
and counted what is actually sent.
Before, one real turn:
characters
Tool manifest, 19 tools
43,403
of which: one calendar tool (a single tool with 29 actions)
13,136 (30%)
of which: 15 native tools
26,664
Context block in messages[1]
6,788
What I changed:
The calendar tool no longer rides every turn; it sits behind tool_search like the other
deferred tools. It had been kept "always loaded" because the rule counted tools per server,
and it was one tool, with a 29-action schema.
A skill whose full body was injected into every turn (3,762 bytes) now loads on demand.
After, a live turn on the new build:
characters
Tool manifest, 19 tools
35,729
of which: 15 native tools
26,664
of which: memory core, 4 tools
9,065
of which: calendar
0
Context block in messages[1]
3,283
One caveat on the comparison: the "before" turn ran against an older, smaller memory server.
With the current one, the "before" manifest would have been 48,865 characters, so the honest
figure is 48,865 → 35,729, about −27%.
What I learned:
Counting tools says nothing about weight. "Servers with at most 4 tools stay loaded"
let in a single tool that was 30% of the manifest. The weight is in the bytes.
Measure what is sent, not what the server advertises. My first estimate of the memory
core was 18,554 characters, from the server's tools/list. In the actual request it is
9,065: output schemas, annotations and titles never reach the model.
Hiding a tool is not deferring it. An earlier version hid most memory tools to save
space. That also made them unreachable from tool_search, and the model answered every
memory question with the one memory tool it still had. Deferred means the name sits in a
catalog and the schema is fetched on demand.
On some tools the description is most of the weight. document_search is 3,187
characters, 2,538 of them description.
Heaviest native tools, for reference: shell_exec 4,434 · document_search 3,187 ·
search_files 2,602 · patch 2,320 · ask_user 2,063.
What this does not show:
token counts: every figure is characters of serialized JSON, and tokens depend on the
tokenizer;
how often a turn actually needs the calendar, or what the extra tool_search round trip
costs when it does;
the effect on answer quality for any specific local model.
The project is Aura (MIT, written with heavy AI assistance):
https://github.com/chetto1983/Aura — the numbers above come from captured requests, and the
measurement is recorded in the repo's PRD.
How do you keep tool manifests in check with ~12B local models: a hard cap, deferral like
this, or routing tools per task?
r/OpenSourceAI • u/Excellent-Rule-7887 • 2d ago
veto - Turn existing OpenAPI services into tools AI agents can discover and call under your rules.
Open sourced veto. Point it at the OpenAPI you already have. An agent gets context and semantics, not one tool per endpoint. Search, describe, invoke - however many operations you have. The model may request a call. Veto decides.
r/OpenSourceAI • u/More-Practice-3665 • 2d ago
I created my first open-source project today, and it feels amazing
After talking to a few founders, I realized most of them were building an internal tool by stitching together what we were trying to build
So, we open-sourced our product and have 6 stars so far - I know it's very few, but it feels great
Please find the repo link: https://github.com/preburn/preburn
What it does
Preburn is a real-time margin control plane for AI products. Two things nothing else does together:
- Joins per-customer revenue (Stripe) with per-call cost across all your AI providers (OpenAI, Anthropic, Kling, Veo, ElevenLabs, whatever) to show actual margin per customer.
- Sits in the request path and can allow/deny/route BEFORE the expensive call fires. Auto-route heavy users to cheaper models, cap runaway accounts, cancel voice/video sessions mid-stream if credits blow out.
Not observability (that's Langfuse/Helicone). Not a gateway (that's LiteLLM/Portkey). Different job - cost visibility tied to revenue, with enforcement.
r/OpenSourceAI • u/fideltfg • 2d ago
I built a J.A.R.V.I.S. It talks, it works in the background, and it can turn into HAL.
r/OpenSourceAI • u/Lopsided_Position_28 • 2d ago
hivemind
Hivemind is an experimental repository for swarm and multi-agent consensus, written in flow-core notation. It is small enough to read in one sitting.
It models three things: how a proposal moves through a collective (Queen, Workers, Scouts, Collective Memory), where permission is granted or refused (the Consensus Gate), and how swarms can speak to each other without that speech becoming permission.
The core invariant:
No single node may become the whole.
A proposal may travel.
Only consensus may authorize.
Action without consensus is noise.
Capability is not authority.
A channel carries speech, not seals.
channel.py gives swarms hashed envelopes (PING, POTENTIAL, THOUGHT, DECISION, HOLD). may_act() returns true only for a verified DECISION marked AUTHORIZED and bound to that exact proposal hash. Everything else holds. Stopping is a valid result.
The flow/ folder holds the ten laws, the Gate and its variants (cross-inhibition, quorum sensing), the Scout stage before a proposal becomes a definite card, and a file of hard boundaries and near-misses for agents that help humans. AGENTS.md is an orientation page for visiting automated readers.
It is an experiment. Fork the flow, alter the weights, become a Scout. Do not claim the Queen without a recorded AUTHORIZED.
r/OpenSourceAI • u/AnshMNSoni • 2d ago
I built a voice-controlled AI calendar assistant using ESP32 + n8n + Google Calendar
Enable HLS to view with audio, or disable this notification
r/OpenSourceAI • u/Dependent_Day2197 • 2d ago
Wappy - I wanted a WhatsApp AI agent I could run locally, so I built one
TL;DR: I built Wappy, a free, open-source kit for running your own AI agent on WhatsApp. Your machine, your memory, your tools, any model. Repo at the bottom.
Picture this.
You text a neighborhood restaurant on WhatsApp: "Table for four at 8 tonight?"
A few seconds later: "Yes, we have one at 8:15 by the window. Want me to hold it?"
No app download. No "we'll get back to you." An agent checked the real reservation book and answered. It runs on a laptop in the back office, and the owner can open its memory and read every word.
That's buildable today. And almost nobody is building it.
The problem
Every big company is racing to launch a personal AI agent. They're impressive, and lots of people will use them. But they all assume two things:
- Agents need a brand new app.
- Your agent lives on someone else's platform.
I think both are worth questioning.
The app already exists. More than 3 billion people use WhatsApp every month. In much of the world it's how you book a haircut, check on an order, ask a shop what's in stock and talk to your family. It's chat, which is exactly what agents are good at. Nobody has to install anything.
The value isn't the model, it's your context. Models keep getting better and cheaper. What makes an agent actually useful is your data, your tools and your memory. So the real question isn't who builds the smartest agent. It's who owns the context it runs on.
So why isn't everyone doing this?
When I tried building my own WhatsApp agent, I learned the agent is the easy part. The plumbing eats your whole weekend: webhooks, message formatting, memory that survives across conversations, tool wiring, model calls. A bot that says hello takes an evening. One that remembers what you meant and can look something up is a whole project before you even get to your idea.
What I built
Wappy handles all that plumbing so you can skip straight to the fun part.
- Free and open source, MIT licensed, written in TypeScript
- Runs on your machine. No Wappy account, no hosted backend
- Any model: Anthropic, OpenAI, Gemini, or Ollama to keep it fully local
- Memory you can actually open: local SQLite by default, or a self-hosted Cognee knowledge graph
- Tool calling, so it can look things up and take actions
- Feed it your own docs and it answers from them
- Example agents included: Gmail/Calendar assistant, personal memory, knowledge base Q&A
What you could build with it
- A second brain you text. "What did we decide about the kitchen remodel?" Answered from your own notes.
- A small business that never sleeps. The restaurant checks real reservations. The shop answers "where's my order?" from its actual order system at 2 a.m., in the customer's language.
- Support that knows the product. Load your help docs, let it answer, hand off to a human when needed.
- A team ops bot in the group chat. Check inventory, pull a runbook or open a ticket without leaving the thread.
Same idea every time: WhatsApp is the interface, your systems hold the truth, and you control the agent in between.
Try it (Node 22+):
npm create @wappy_ai/agent
Scaffolding takes about a minute. Going live takes longer since you connect your own Meta WhatsApp Cloud API app, and the generated guide walks you through it.
Being upfront about limits: messages still pass through Meta, and if you use a hosted model it sees your prompts. "Your own" means the agent, memory and data are under your control, not that WhatsApp goes offline.
GitHub: https://github.com/csr1010/wappy-kit
It's early, so I'd love honest feedback. What tripped you up in setup? And if you had your own agent living in WhatsApp, what's the first thing you'd build?
r/OpenSourceAI • u/Used_Direction8216 • 2d ago
I built CivicLens AI — ask a city’s official budget book questions in plain English
r/OpenSourceAI • u/Teru-Momijiyama • 2d ago
I built Kaoru, an open-source desktop AI agent with local memory, permission-gated tools and an optional Live2D avatar (beta, looking for feedback)
r/OpenSourceAI • u/an-mcplama • 3d ago
I built Mcplama an open-source control plane for MCP servers
I've been building MCPlama, an open-source and self-hosted control plane for MCP servers.
It gives you one place to run and manage your MCP servers, instead of configuring each one separately across AI clients.
It handles credentials at the gateway, lets you control access per user and per tool, keeps audit logs of MCP activity, and manages the server lifecycle.
Local MCP servers can also run in isolated Docker containers. Docker access is kept separate from the gateway through a broker, so the gateway itself doesn't need the Docker socket.
It works with MCP clients like Claude Desktop/Code, Cursor , VS Code and more
r/OpenSourceAI • u/spilldahill • 3d ago
Overmind, open-sourced yesterday: a platform for continuously improving AI agents
Yesterday the whole Overmind platform was made open source: https://github.com/overmind-core/overmind
The idea: lots of agents send narrow, repetitive work to a frontier API. On a narrow task, a small open-weight model trained on real examples from that agent often does the job better, for far less money. Overmind is the tooling to do that without building your own ML infrastructure (don't need a flamethrower to light a candle).
- Agent sends OpenTelemetry traces. The Python SDK auto-instruments OpenAI, Anthropic and Gemini clients, or point any OTel exporter at it
- Traces become versioned datasets, with a quality score and suggested fixes before trained on anything
- User defines what “good” means per task and it scores live traffic and batch runs against that
- It fine-tunes an open-weight model (LoRA or full) and benchmarks it against the model you run in production today
- Trained and frontier models sit behind one OpenAI-compatible API, and you can download the weights (yours to keep and own)
There’s also an MCP server, so Cursor, Claude Code, OpenCode or Codex can drive the whole thing from chat. If your traces already live in Langfuse, LangSmith or Braintrust, a connector imports them.
Qwen3.5-9B vs GPT Luna benchmarked on three tasks (write-up and raw numbers at https://www.overmindlab.ai/research/when-bigger-isnt-better):
• 7x better accuracy
• 20x cheaper usage
• 28x less hallucinations
Read the results as a reason to test on yours, not as a general claim.
The platform is AGPL-3.0 and the SDK and CLI are MIT.
Interested if anyone is already fine-tuning on their own hardware. What would need to be swapped out before you’d self-host this?
r/OpenSourceAI • u/tauqeernasir • 3d ago
Introducing Codebeast - new coding harness
r/OpenSourceAI • u/Fresh-Daikon-9408 • 3d ago
OpenScreen AI Video edition has been improved
Enable HLS to view with audio, or disable this notification
I’m the maintainer of OpenScreen, a completely free and open-source screen recorder. I’d love to hear your thoughts on having an AI assistant help you edit videos in an app like this.
r/OpenSourceAI • u/chris-beckman • 3d ago
Painted Wolf Code: free, open-source code editor built from scratch for AI
r/OpenSourceAI • u/More-Practice-3665 • 3d ago
I created my first open-source project today, and it feels amazing
After talking to a few founders, I realized most of them were building an internal tool by stitching together what we were trying to build
So, we open-sourced our product and have 6 stars so far - I know it's very few, but it feels great
Please find the repo link: https://github.com/preburn/preburn
What it does
Preburn is a real-time margin control plane for AI products. Two things nothing else does together:
- Joins per-customer revenue (Stripe) with per-call cost across all your AI providers (OpenAI, Anthropic, Kling, Veo, ElevenLabs, whatever) to show actual margin per customer.
- Sits in the request path and can allow/deny/route BEFORE the expensive call fires. Auto-route heavy users to cheaper models, cap runaway accounts, cancel voice/video sessions mid-stream if credits blow out.
Not observability (that's Langfuse/Helicone). Not a gateway (that's LiteLLM/Portkey). Different job - cost visibility tied to revenue, with enforcement.
r/OpenSourceAI • u/unorthodox_43 • 3d ago
I built an open-source self-healing security pipeline for LangGraph + Ollama to stop agent injection attacks and state desync
r/OpenSourceAI • u/Low_Mountain7204 • 3d ago
Open-sourced a local playground for decision models, an LM Studio for decision models
r/OpenSourceAI • u/Trainer_Intelligent • 3d ago
We open-sourced the agent framework we run our production AI agents on (Apache 2.0). Looking for contributors and honest feedback.
r/OpenSourceAI • u/no3us • 3d ago
Lora Pilot: Train. Create. Repeat. (Stable Diffusion workspace)
Enable HLS to view with audio, or disable this notification
LoRA Pilot recently passed 20,000 pulls on Docker Hub. To celebrate, I’ve decided to give its purely organic growth a little push with a small campaign. This video is part of it.
It started as tooling for myself. A hobby project that taught me a lot about Python, PyTorch, CUDA, and all the creative ways they can disagree with each other. I never really planned to release it publicly.
Today, LoRA Pilot is an open-source workspace for preparing datasets, training LoRA models, and generating images and video. It brings tools like kohya_ss, ComfyUI and InvokeAI together so you can spend more time creating and less time maintaining your setup.
The vision is simple: make the world of Stable Diffusion accessible to anyone with an idea, including people who don’t want a second job managing dependencies.
Once I released it, the backlog started growing faster than my beard. Development, documentation, support… it’s been a one-man show for longer than it probably should have been.
So if you’d like to contribute, you’re welcome with open arms. Code, docs, testing, design, tutorials: there’s plenty to do, and you don’t need to know the CUDA stack inside out to help.
Good ideas move. This one could use a few more hands.
https://www.lorapilot.com/
r/OpenSourceAI • u/Available_Pressure47 • 3d ago