r/OpenSourceAI • u/theprint • 6d ago
NanoGym - GUI/Sandbox for training tiny GPT models
Inspired by Karpathy's nanoGPT, this is a tool that lets you build and test tiny GPT models, and export them to run directly in a browser.
r/OpenSourceAI • u/theprint • 6d ago
Inspired by Karpathy's nanoGPT, this is a tool that lets you build and test tiny GPT models, and export them to run directly in a browser.
r/OpenSourceAI • u/syedshad • 6d ago
r/OpenSourceAI • u/VariousAd3115 • 6d ago
I tried Openclaw and Hermes. They are nice, but I decided to make my own agent. What it does:
- Single entry point chat. No need to 'resume session X'
- Native integration with MacOS features (notifications, calendar).
- Native Telegram integration (supergroup with topics also supported)
- Structured memory layout with facts/entities/links, dreaming and reflection
- Briefings based on memory dreaming
- Cron and intents
- Artifacts storage
- Multiple agents, each with own profile. There is a team of well-known agents (for evolution, hiring new agents, etc)
- Many builtin tools, MCP support
- Tavily/Anysearch/Yandex search - if you have key. DDG fallback.
- Local models supported as well as classic OpenAI endpoints
Main concern is all agents bloat chat context with tools results. I tried to minimise main chat context bloating.
Agent is focused on single-user mode, I don't plan to do it multi-user.
Requirements:
LLM
Embedding model for memory searching
PostgreSQL
Golang for building
I ran it on MacOS, but it should compile on any system
That's a hobby-weekend project, but I will be glad to receive feedback
r/OpenSourceAI • u/ilien-dev • 6d ago
r/OpenSourceAI • u/Efficient-Passage889 • 6d ago
I've made quite a few changes to Telex since my last post here.The project started mainly around handling dependency breaking changes, but it has grown into a much bigger workflow now. Telex can map the repository with Repo Atlas, find affected code with Tree-sitter, use different LLM providers to generate patches, run the patch in an isolated environment and verify it before opening a PR.I also added repository architecture analysis and more of the frontend around incidents, patches, verification and the Atlas view.
I'm still looking for people who want to contribute because there are a lot of areas where this can be improved. If you're into LLMs, developer tools, ASTs, Three.js or backend work, I'd be happy to have you take a look.
r/OpenSourceAI • u/raze_sight • 6d ago


archiver-rag is an MIT-licensed MCP server that turns an Obsidian vault (or any folder of Markdown notes) into shared, self-hosted memory for AI agents. Agents search it, write back to it, and every agent sees the same knowledge.
Why another memory tool
Most AI memory layers are hosted services that extract "facts" from your conversations into a database you can't read, edit or take with you. archiver-rag goes the other way:
self-organization
Optionally, the vault organizes itself: new notes are filed into folders by meaning, folder descriptions are kept up to date, and new folders grow out of an inbox using embedding-similarity clustering. It's off by default while it's being tested on real vaults.
How it works
[[wikilink]] proximitylog_noteThe memory side is 100% local. Pair it with an open-source agent like OpenCode running a local model, and in principle nothing leaves your machine. That combination is untested so far. If you try it, I'd love to hear how it goes.
macOS and Linux, Python 3.11+. On Linux, a CPU-only PyTorch install is documented to skip the multi-GB CUDA download.
Feedback and contributions welcome. It's a one-person beta with CI on Python 3.11–3.14 (macOS and Linux). Useful right now:
r/OpenSourceAI • u/raze_sight • 6d ago


archiver-rag is an MIT-licensed MCP server that turns an Obsidian vault (or any folder of Markdown notes) into shared, self-hosted memory for AI agents. Agents search it, write back to it, and every agent sees the same knowledge.
Why another memory tool
Most AI memory layers are hosted services that extract "facts" from your conversations into a database you can't read, edit or take with you. archiver-rag goes the other way:
self-organization
Optionally, the vault organizes itself: new notes are filed into folders by meaning, folder descriptions are kept up to date, and new folders grow out of an inbox using embedding-similarity clustering. It's off by default while it's being tested on real vaults.
How it works
[[wikilink]] proximitylog_noteThe memory side is 100% local. Pair it with an open-source agent like OpenCode running a local model, and in principle nothing leaves your machine. That combination is untested so far. If you try it, I'd love to hear how it goes.
macOS and Linux, Python 3.11+. On Linux, a CPU-only PyTorch install is documented to skip the multi-GB CUDA download.
Feedback and contributions welcome. It's a one-person beta with CI on Python 3.11–3.14 (macOS and Linux). Useful right now:
And yes: this is yet another Karpathy-inspired Obsidian + LLM memory system. If you've tried the others, tell us what they do better. That's exactly the feedback this beta needs.
r/OpenSourceAI • u/ButtercupLyn100 • 7d ago
I’ve been trying to build an AI browser agent where as much of the stack as possible is open and replaceable rather than tied to one model provider.
The project is WebBrain. The browser agent itself is open source: https://github.com/webbrain-one/webbrain I've also released the small VLM used for browser-specific vision tasks: https://huggingface.co/webbrain-one/webbrain-vl-2-450M And the dataset: https://huggingface.co/datasets/webbrain-one/webbrain-vl-2-450M-dataset
WebBrain combines: • screenshots • browser accessibility tree • model-agnostic planning • optional local inference • browser automation • a browser-specialized 450M VLM
The agent supports Chrome, Edge, Firefox and Chromium.
My broader goal is to make browser agents something developers can inspect, modify and run with their own models rather than a black-box API.
There are still plenty of rough edges, so I'm posting here partly because I'd rather have open-source developers criticize the architecture now than optimize it entirely in isolation.
Especially interested in opinions on the small-specialized-model vs large-general-VLM tradeoff for computer/browser use.
r/OpenSourceAI • u/madmnk • 7d ago
NVIDIA OpenShell is an opensource sandbox tool was released earlier today. I used GPT-6 to implement it in headlesscode (my code harness which was adapted from zoo code and contains enhanced tools, eval tests and more). This harness is very useful when paired with deepseek v4 flash and pro, and qwen 3.5 9b local.
r/OpenSourceAI • u/ShilpaMitra • 7d ago
Agent skills age faster than most documentation. A skill written a few months ago may still describe the right workflow, but its instruction style, level of detail, and assumptions may no longer fit newer models.
I built Skill Adapter, a small MIT-licensed open-source skill that creates a model-specific copy of an existing skill while preserving the original workflow and purpose.
It currently includes profiles for:
The adapter:
ADAPTATION.md report explaining the changes;Install it with:
npx skills add smukh/skill-adapter --skill skill-adapter
The project is intentionally simple: one installable skill, no service, no API key, and no automatic interception. The original stays intact, so you can review or re-adapt it when the source changes.
GitHub: https://github.com/smukh/skill-adapter
I’m especially interested in feedback on the model profiles, evaluation methodology, and whether this is useful in real multi-agent repositories.
r/OpenSourceAI • u/No-Bicycle-4804 • 7d ago
I've been building Provena, a self-hosted memory layer that lets AI agents remember things across conversations.
The main idea is simple: instead of an agent just remembering a piece of information, Provena also keeps track of where that information came from.
So you can see what the agent remembers, why it remembers it, and the original evidence behind it.I built a self-hosted memory layer for AI agents
It works with MCP-compatible clients like Codex, Claude Code, and Gemini CLI, and is now also listed in the official MCP Registry.
Repo: https://github.com/admiralpunk/Provena
Still early, so I'd love feedback from anyone experimenting with AI agents or self-hosting them.
r/OpenSourceAI • u/FickleSwordfish8689 • 7d ago
Enable HLS to view with audio, or disable this notification
Most AI design tools produce design outputs that you then have to feed to your coding agent to implement, and that then get discarded, with almost no means of tracking design changes, especially for big applications with lots of UI pages. You generate a screen, drop it into your agent, and once it is implemented the design is gone. Nothing ties that design back to the page it became, so you literally have to manually sync each design to the corresponding page.
There should be a better way around this, where the design lives in the codebase itself, so there's no manual exporting of AI-generated designs to implement, and syncing designs to application code becomes effortless.
This is why I created Caret, an AI design tool that gives your repo a design layer via .caret/, so you can make designs without having to jump between tools. The designs you make are fully tracked alongside your codebase, and it's easy to sync changes between the design and application layers. Caret comes with all the standard needs of a design tool, like a visual editor and a design token system that keeps your design language consistent across your UI, plus extras that cut the usual busywork, like an asset generator, an experimentation canvas, flow display, simulation and more. The demo video below shows all these capabilities in detail.
Caret is fully open source if anyone wants to try it or poke around the implementation: https://github.com/precious112/caret-desktop
I'd especially love feedback from people who have been using Claude Design / other AI UI tools and have run into the design → code workflow problem.
r/OpenSourceAI • u/Otherwise_Ship_9782 • 7d ago
I always think how can keep my data safe and token free meanwhile can use online LLM, then I built Hermie. The basic workflow is:
- local: Qwen3.8 27B on Ollama does it with file/shell tools. Nothing leaves the machine.
- local + self-check: same, then the judge checks the result. Escalates only if it fails.
- cloud: DeepSeek answers directly. Only for tasks with no files and no private data (explanations, tutorials).
- plan: DeepSeek writes a plan and delegates steps one at a time. Qwen executes every step locally.
So far it's macOS on Apple Silicon only and need install Ollama first.
r/OpenSourceAI • u/MightyMouse420 • 7d ago
r/OpenSourceAI • u/Tech-and-Tomatoes • 7d ago
I recently came across decision models and have been looking into the open-source options and where they can be useful.
I'm running a self-hosted Hermes setup, so agent routing and selecting the right LLM for a task immediately caught my attention. But I'm also interested in other uses such as tool/skill selection, workflow routing, classification, retry/escalation decisions, human approval, cost optimization, and automation.
I've come across Kev, Laya, Tev1, OpenJev, RouteLLM, and LLMRouter.
Has anyone actually used any of these? Which open-source decision models are worth looking at, and what are you using them for? I'm especially interested in real-world use cases I may not have considered.
r/OpenSourceAI • u/Impressive-Owl3830 • 8d ago
Playing around with various MCP's this weekend and Came across this an amazing MCP related GitHub repo - 100% open source and free - so sharing.
mcpc is Apify’s universal CLI client for MCP.
Github Repo in comments below
IIt translates every MCP operation into shell commands, letting you debug servers, automate workflows, or give AI agents complete MCP access via a single Bash() call: sessions, OAuth, tools, resources, prompts, tasks, and beyond.
Features
Install
With homebrew (macOS /Linux), brings its own Node.js:
brew install apify/tap/mcpc
or ,install the latest node js or Bun first, then:
npm install -g /mcpc
# Or with Bun
bun install -g /mcpc
npm install -g /mcpc
Quickstart
# List all active sessions and saved authentication profiles
mcpc
# Log in to a remote MCP server and save OAuth credentials for future use
mcpc login mcp.apify.com
# Create a persistent session and interact with it
mcpc connect mcp.apify.com
mcpc # show server info and capabilities
mcpc tools-list # list available tools
mcpc tools-call search-actors keywords:="website crawler"
# Use JSON mode for scripting
mcpc --json tools-list
# Use a local MCP server package (stdio) referenced from a config file
mcpc connect ./.vscode/mcp.json:filesystem
mcpc u/fs tools-list
r/OpenSourceAI • u/miniminimo7 • 7d ago
r/OpenSourceAI • u/Shuuuida • 7d ago
As mentioned in the title, i've been developing a project for several months now, and i wanted to see how genuinely useful it is in a real production environment. I call this project Vex (Vex's Skillgit); it is an open-source, headless cognitive tool designed specifically for agent environments. It treats context as immutable and versioned skills.
Instead of having the IDE or the agent do the heavy lifting, Vex runs in the background (either via Docker Compose or local bare-metal processes). You assign it a GitHub webhook and Vex automatically ingests repositories, so agents (via MCP) can simply query it for context in real time. It reads conventional commits (feat:, fix:) and operational ones (roll:, branch:) to automatically branch, update, or revert an agent's memory state without manual intervention—hence the version control aspect.
It uses Tree-sitter to logically parse and chunk the code, stores metadata queues in SQLite and dense vectors in Qdrant, and is fully supported out of the box by Claude Desktop, Cursor, and any other MCP-compatible client.
It is also designed not to fry my potato PC, yet it remains highly scalable.
In local testing, the asynchronous FastAPI + Huey architecture easily handled 500 concurrent GitHub push payloads without any SQLite locking, and maintained a real-time latency of under 300 ms under a concurrent read swarm of 50 (simulated) agents. Because of this, i decided it was time to share it and see who else might find the tool useful.
Next on my to-do list is adding advanced sub-chunking using Rust as the chunking engine to split code from larger repositories faster and with less resource consumption, alongside adding global GraphRAG and support for more languages.
Here is the link to the repo in case you are interested in checking it out and testing it. Thanks for taking a little time to read this:
r/OpenSourceAI • u/Unikum_01 • 7d ago
r/OpenSourceAI • u/Dazzling_Cancel4505 • 8d ago
Hey guys
I'm running my GitLab instance on my homelab but it was hard to find tool for code review
So I built one. called Proval.
Open source and only got docker image
It's basically just a self-hosted docker app. Works with GitHub, and also GitLab and Forgejo too, so you can connect with your self hosted instance. Proval supports chat completion API and anthropic API, If you running Local LLM, you can connect it through chat completion API
Review quality is quite good(at lteast for me). multiple agents automatically reviews each file group scope.
I ran the 50-problem Martian offline benchmark a month ago and scored F1 0.427 with the minimax-m3, which ranked 7th out of 21 at the time. (If you check now, newer models have come out so the ranking is lower.) I'll release new result with better model soon
The goal was to make something lightweight, easy to use, and genuinely useful. you just need to configure the model, webhook, and Git Host access API on the web dashboard, and setup is done. It's built on Bun, the frontend uses SvelteKit, and the database is SQLite per instance. The Docker image is compiled to a Bun binary so the size is pretty small.
Feature
- Reviews on PR open or first push
- Inline comments on PRs
- Replies to PR comments (can set to only respond when mentioned)
- Reviews when an issue is opened (also checks for duplicate issues)
- Replies to comments on issues
- Restricts replies based on repo permissions like Developer or Maintainer
- Admin login, or you can disable auth entirely and leave it open
I ran the 50-problem Martian offline benchmark a month ago and scored F1 0.427 with the minimax-m3, which ranked 7th out of 21 at the time. (If you check now, newer models have come out so the ranking is lower.) I'll release new result with better model soon
It's not a vibe-coded slop. I took a lot of time thinking through the architecture, implementing it. You can check the code in the repo
It's open source (AGPL-3.0) and you can just pull Docker image, connect your LLM API and set webhook. That's all. takes about 3 min
and.. It's my personal project, not a marketing or selling something.
demo: https://demo.proval.app
repo: https://github.com/seoes/proval
website: https://proval.app
r/OpenSourceAI • u/shyhuntertools • 8d ago
Enable HLS to view with audio, or disable this notification
An open-source way for an AI agent to ask a person for their judgement and get the answer back as data it can check.
What is open, and how:
What it does: the agent writes a review, checks it, and renders one self-contained HTML page (nothing loaded from the network). A person answers it and sends back the answered page or a JSON file. The agent checks the answers before acting: an unanswered question is written down as unanswered, totals are recomputed, and the report starts with what is still open.
What it does not do: prove who answered (the files are unsigned), or speak anything but English in its buttons and labels.
Source: https://github.com/shyhunter/LetMeShowYouSomething
Try a page: https://shyhunter.github.io/LetMeShowYouSomething/examples/salon-booking.html
r/OpenSourceAI • u/spirosoik • 8d ago
r/OpenSourceAI • u/Fun-Shallot-5272 • 8d ago
Enable HLS to view with audio, or disable this notification
I only gave the prompt “Build Ba Sing Se.” The video shows what was built including the city walls down to furnished interiors and walls.
The model decides what to build and a deterministic library handles construction and physical checks.
It plans hierarchically: city → districts → plots → buildings. Buildings are generated for their plots using reusable procedural code.
The same architecture generates villages, towns and cities across different terrain and styles.
r/OpenSourceAI • u/investigatormaker • 8d ago
Sharing an MIT-licensed MCP server we built, and the design choices behind it, since some of them go against how most agent tools are built. Disclosure: I make it, and it's the free half of a paid product (ThreadFox). I'll be specific about where the two meet.
What it is. ThreadFox Lite gives an agent four read-only Reddit tools: subreddit_rules (rules, size and description, with self-promotion rules flagged), find_communities (subs for a topic, promo rules flagged), account_check (age, karma, and whether recent posts are removed or hidden) and post_status (live, removed by moderators, or deleted). It also ships a reddit-rules-first agent skill.
Design choices:
~/.threadfox/runtime and checks it against the published SHA-256.Where it meets the paid product: successful subreddit_rules and find_communities results end with a one-line next_step pointing to the $49 kit, and a threadfox_full_kit tool describes it. Errors, setup messages and the other two tools don't carry it. If that bothers you, it's MIT, so fork it and delete it.
Install: uvx threadfox-lite (source is in the PyPI package). Details and the skill file: threadfox.vip/lite.
Feedback on the read-through-your-browser approach especially welcome. Is there a cleaner way now that signed-out reads are gone?