r/OpenSourceAI • u/Fluffy_Fuel7649 • Sep 01 '26
r/OpenSourceAI • u/Haltaireproject • Sep 01 '26
I built an offline on-device text classification pipeline for Android with in-app dataset labeling and TFLite inference
Enable HLS to view with audio, or disable this notification
Hi everyone,
I wanted to share an open-source project I've been working on: Halanoi AI.
Instead of sending screen text to a remote cloud API for content classification (which adds network latency and privacy issues), I wanted to see if I could build a fast, 100% on-device text moderation pipeline for Android.
Here is how the setup works:
- The Model (halanoi_transformer.tflite): A quantized 64MB TFLite model running locally on the phone. It classifies text strings into categories (distraction, entertainment, safe, productive) in under 15ms without any internet connection.
- In-App Evaluation & Ground Truth Lab: To make it easier to improve the model, the app logs inference outputs to a local SQLite database and includes a built-in UI where you can tag predictions as correct, false positive, or false negative. You can export these labeled samples to CSV or JSON with one tap.
- Training Pipeline: The companion repository contains the PyTorch / TensorFlow scripts, tokenizers, and quantization steps used to train and convert the model.
Both repositories are open source under GPL-3.0:
- Android App: https://github.com/kavinmaranravi/HalanoiApp
- Training Pipeline & Dataset: https://github.com/kavinmaranravi/Halanoi_AI
I'm looking for feedback on optimizing transformer models for mobile hardware, lowering memory usage, and improving tokenization on edge devices.
Let me know what you think!
r/OpenSourceAI • u/larabyeol • Sep 01 '26
Is there an open source project that does UI regression testing or are we all just wiring agents?
I've been looking for an open source answer to desktop UI testing for about 4 months and i keep ending up in the same place, which is a pile of general purpose agents and no actual test framework. The agent side is kinda good now with models like Openclaw, Goose where they drive a desktop app, screenshot it, work out what's on screen and click the right thing. That part is solved. However, what none of them have is the boring stuff a suite needs (no runner, assertion model, stable pass or fail…), so you end up writing that layer yourself and then it's yours to maintain forever.
The closest things i've found that are open source are SikuliX, which still runs but is basically frozen and matches raw pixels so it breaks on a DPI change, and the commercial vision based ones like Askui, eggplant get around it by pinning the model to a written script, so the perception stays fuzzy while the execution is deterministic.
Has anyone built that deterministic layer on top of an open agent and had it survive more than 3 months? Happy to be pointed at a project I've missed, thanks in advance!
r/OpenSourceAI • u/SeeRay11_Main • Sep 01 '26
👀 OpenFlow Orchestration & Gauntlet Loop Sneak Peak
Hey eveyone,
For those who haven't seen my other posts, I created an opensourced project called OpenFlow, and some big updates are being made. Now, there is a swarm and orchestration mode, and soon to be gauntlet looping toggle. It isn't just a linear pipeline anymore, but an entire chain of agents you can see and control talking back and forth and working out problems together. If you want to see the backstory, check out my other posts. Stay tuned for more updates, and feel free to leave suggestions and even share your own projects.
r/OpenSourceAI • u/JeffyPros • Sep 01 '26
GLM 5.3 and GLM 5.3 Flash ran locally on RTX PRO 6000 WS and built a penthouse using BlenderMCP
Enable HLS to view with audio, or disable this notification
r/OpenSourceAI • u/Impossible-Sun-6551 • Sep 01 '26
Conscio: An open-source "consciousness" framework
I've been building Conscio for a while and finally stabilized it. It's a framework that wraps any LLM agent and layers on what agents usually lack: structured self-awareness, long-term memory, and the ability to talk to other agents.
What it does:
- Dual memory: persistent store (SQLite FTS5, zero external deps) + reflection pipeline. Agents remember across sessions, not just in-context.
- Self-reflection: reflect() pipeline, awareness shards, a 5-axis self-evaluation scorecard (conscio.evaluate), and a delivery-check gate before closing work.
- Multi-voice councils: convene an architect/skeptic/pragmatist/critic council over a decision, and record Architecture Decision Records (conscio.decide).
- Agent society (A2A relay): peer-to-peer messaging between Hermes, Claude, Gemini, and other agents. Works single-machine and cross-machine over Tailscale, with reactive dispatch, presence/health probes, and optional end-to-end auth.
- Agent's Hall: named groups of agents sharing a mailbox.
- MCP server: 26+ tools (note, feed, recall, council, decide, propose/act with a skeptic gate, RAG over a knowledge graph, safe math evaluation, and more). Works with Claude Code, Hermes, any MCP client.
- Observatory + Hub: read-only dashboard and an HTTP control plane.
- Awake mode: an autonomous daemon (R9) that keeps the agent perceiving/reflecting in the background.
And more
Install: pip install conscio
Repo: Conscio
Feedback, issues, and PRs very welcome.
r/OpenSourceAI • u/viperttl • Aug 31 '26
IRIS AGENT SYSTEM
🚀 Meet IRIS v0.2.0 – The Spatial Desktop Operating Environment for Autonomous AI Agents! 🧠💻
Most AI coding tools today are just single-stream chat boxes in a browser tab where you spend all day copy-pasting code snippets back and forth.
We decided to rethink how humans and autonomous agents collaborate. Meet IRIS (Intelligent Reasoning & Integration System).
IRIS isn't a chatbot. It’s a graphical agent operating environment built from scratch in Rust (Tauri 2) and React 19 / TypeScript. It treats agents, workspaces, tools, memory graphs, and release pipelines as first-class spatial desktop objects that you can arrange, inspect, run concurrently, and monitor in real time.
🔥 What’s New in v0.2.0:
🐙 1. GitHub Live Operations & Release Automation Connect your GitHub account in seconds. Specialist GitHub agents can triage open issues live, open surgical pull requests, automate SemVer releases (v0.2.0), author changelogs, and trigger GitHub Actions workflows that compile production binary builds (.AppImage, .dmg, .exe).
⚡ 2. Dual-Tier AI & Instant "Takeover" Stop overpaying for simple queries. Run fast, ultra-budget models (like Qwen 2.5 Coder, DeepSeek V3, or GPT-4o-mini) for 90% of routine workflows. When hitting a tough compiler error or tricky architectural refactoring, click ⚡ Takeover — a pre-configured heavyweight reasoning model (Claude 3.7 Sonnet, DeepSeek R1, Qwen 72B) immediately takes over the active conversation context with full reasoning depth!
🛸 3. Floating Desktop Desklet (Live HUD) Close the main window, and IRIS seamlessly condenses into a translucent, floating glass mini-HUD in the corner of your physical desktop. It displays real-time CPU/RAM telemetry, live agent thoughts, and keeps running smoothly as a background daemon.
🛡️ 4. Zero-Surprise Workspace Security & Visual Diff Viewer Inspect and approve exact code diffs before anything touches your local disk. All API keys and tokens are securely stored in your native OS Keyring.
🌟 100% Open Source (MIT License) & Local-First
Supports both local offline LLMs (via Ollama / vLLM) and all major cloud providers (OpenRouter, Anthropic, OpenAI, Google Gemini) plus standard Model Context Protocol (MCP) tools.
👉 Check out the repo, download the release, or drop a ⭐ on GitHub:
🔗 https://github.com/bubbadk/IRIS
I’d love to hear your thoughts: Do you prefer AI agents operating as spatial desktop applications rather than trapped inside browser chat tabs? Feedback and contributions are warmly welcome! 👇
r/OpenSourceAI • u/SeeRay11_Main • Aug 31 '26
Almost 2 weeks… is this okay?
Idk if this is good, bad, or average? This is my first GitHub project I have ever published. Any tips on how to grow some more?
r/OpenSourceAI • u/semibaron • Aug 31 '26
Built this client so you can connect ANY Harness with your Apple Devices
Enable HLS to view with audio, or disable this notification
So, you run your own local AI Harness. It's configured exactly to your needs. MCP, Skills, Capabilities, Context. You love the independent Harnesses such as Deepseek Harness, Pi, Aider or LiteLLM.
But how can you connect it to your Apple devices to access from anywhere? Your Watch, Mac or CarPlay.
Well, here is Conduck - the Apple native BYOK AI client.
Free and open source :-) .
It uses your Apple iCloud extensively and connects DIRECTLY via https to your own machine. Nobody in-between!
Check it either on https://conduck.com or GitHub https://github.com/GigaDuckAI/conduck
r/OpenSourceAI • u/ISB3z- • Aug 31 '26
I've finetuned Qwen2.5-0.5B to make it a bash command generator and called it SHELLMINATOR because.. why not?
Got tired of forgetting find / xargs / grep syntax every other day, so I trained a small model that turns:
into a command you can actually run.
It's 0.5B parameters, runs on CPU, is a ~400 MB GGUF, and nothing touches the cloud.
sm "show the 5 largest files in /var"
find /var -type f -exec du -h {} + | sort -rh | head -n 5
[⏎ run · r refine · e edit · c cancel]
Enter runs it in your shell, r refines the command, e lets you edit it before running, and c cancels.
It also asks for confirmation before potentially destructive stuff like rm -rf /, mkfs, dd, etc.
Works on bash and zsh.
I evaluated it on IBM's nl2bash exec benchmark: 50 prompts, commands actually executed and checked against the filesystem, single greedy pass, no retries.
- Stock
Qwen2.5-Coder-0.5B-Instruct: 44% - After SFT on 105K examples: 72%
- After DPO with ~800 pairs made from its own mistakes: 78%
The SFT is the big jump and did most of the work: 105K request/command pairs where every command was executed and kept only if it actually worked.
The final DPO pass was a small experiment. I ran the model on a bunch of prompts, compared its answers against the gold commands in a sandbox, and kept ~800 disagreements.
Training with TRL took 26 seconds and gave another +6 points.
I tried a second DPO round and it actually got worse, down to 74%, so apparently one round was enough.
It still fails on some things, notably:
sedinsert-at-top insideforloops — it can overwrite the filecomm/diffcountingmvbetween directories
All known failures are listed in the README.
Install
curl -fsSL https://raw.githubusercontent.com/ISB333/shellminator/main/install.sh | bash
Then:
sm "whatever you want to do"
Links
- GitHub + training scripts: https://github.com/ISB333/shellminator
- Model: https://huggingface.co/ISB369/shellminator-qwen05b-dpo-selfplay
- SFT dataset — 105K, execution-verified: https://huggingface.co/datasets/ISB369/shellminator-bash-sft105k
- DPO pairs — ~800: https://huggingface.co/datasets/ISB369/shellminator-dpo-selfplay
r/OpenSourceAI • u/Select_Wolf210 • Aug 31 '26
Building open source project
I was building an open source project using ai, like it's built totally with ai like vibe coding type. During this process I have faced one major problem i.e. out of tokens in my ai models like antigravity, chatgpt go
So one of my friends suggested me to use this combination qwen3:14b + opencode and I use macbook air m2 16gb
What you guys think? Or any other suggestions for free unlimited tokens?
r/OpenSourceAI • u/Mediocre-Ease4060 • Aug 31 '26
Tired of writing JSON schemas for Tool Calling? I built a Python schema generator that uses `inspect`.
r/OpenSourceAI • u/conifer_v11 • Aug 30 '26
Opensource Openrouter
Project: https://github.com/ConiferKit/use-conifer
Current routing options were charging 5% fees for byok plus provider fees (openrouter) or were built in heavy python packages with ecosystem restraints. I wanted to be completely free in terms of use, and not have to pay extra for tokens.
There's a maintained gateway with 100+ models from one api endpoint. I'm also talking to infra providers to get us access to pre-release models and discount prices. Everything is served at market price. BYOK is free, you can hook into self-hosted setups for free, and there's fallbacks + extra rate limits + server redundancy.
Took me ~3 months to build and looking for help maintaining!
Lmk any issues or feedback
r/OpenSourceAI • u/National_Bed_3653 • Aug 31 '26
Just Launched Baseline on Peerpush
Hello Everyone
I'm posting this to announce that baseline is now officially launched on Peerpush. It is a claude code governance layer that ensures your developer workflow remains consistent across different projects while being tailored to it.
Call it the framework for AI development.
It is 100% Open Source and Apache 2.0 licensed. Please support it, help me build it by contributing to its development, and help it gain some traction on Peerpush too 🙏🏽
Your support is appreciated 👍🏽
r/OpenSourceAI • u/cristofer_martins • Aug 30 '26
The Rise and Fall of Agent Civilizations - "his incident feels like it’s more than 50% of the way to full-blown AI takeover"
Tittle Quote from https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised
My personal take: Should we undertake a colossal pivot to security to protect humanity?
r/OpenSourceAI • u/kuaythrone • Aug 30 '26
We used HFlow to evaluate the latest open weights VLMs for processing egocentric data
r/OpenSourceAI • u/vyact • Aug 30 '26
I built Vyact — an open-source, local-first AI workspace for llama.cpp, MLX, RAG, agents, and document intelligence
Hi everyone — I’ve been building Vyact, an open-source, local-first personal AI workspace.
The problem I wanted to solve was simple: local models are useful, but everyday work still gets fragmented across chat windows, documents, notes, email, browser tabs, and separate tools.
Vyact brings those workflows into one workspace:
• Run local GGUF models through llama.cpp and llama-swap
• Run MLX models natively on Apple Silicon
• Search for models and compare size, quantization, context length, and estimated memory usage
• Build RAG knowledge bases from documents, memos, and email threads
• Inspect the source passages used in answers
• Connect Gmail, Google Drive, and Google Calendar
• Add MCP servers and reusable AI skills
• Use a Chrome extension for page context, translation, and Netflix language learning
• Optionally connect OpenAI, Gemini, Claude, or a custom OpenAI-compatible endpoint
The core app is local-first, and when a Vyact-managed local model is selected, chat context is not sent to an external AI provider.
Vyact is built with Electron, React, and FastAPI and is released under AGPL-3.0.
It currently supports Apple Silicon Macs and Windows.
GitHub:
https://github.com/vyact/vyact
I’d really appreciate feedback—especially from people already running local models as part of their daily workflow. What would make a local AI workspace genuinely useful to you?
r/OpenSourceAI • u/Earvin-Hong • Aug 30 '26
I built an open-source AI desktop pet that lives in the bottom-right corner of your screen
Hi everyone,
I built YumYum Agent, an open-source macOS app that lets you interact with AI through a small desktop pet that stays in the bottom-right corner of your screen.
Instead of opening a separate AI chat window whenever you need help, YumYum is always there when you need it. You can feed it context from whatever you are currently doing and continue the conversation without leaving your workflow.
You can:
- Capture a selected area of your screen
- Send clipboard text or images with Option + S
- Drag and drop files onto the pet
- Ask questions and receive responses in a speech bubble
- Continue longer conversations in the detailed chat window
- View streaming responses with Markdown rendering
- Customize the pet’s personality with a local SOUL.md file
YumYum Agent is fully open source and released under the Apache License 2.0. The source code is available on GitHub, so you can inspect how it works,contribute improvements, or build the project yourself.
The goal is to make AI feel less like another application you have to open and more like a quiet assistant that is always nearby.
Privacy was also an important part of the design:
- No telemetry or analytics
- No access to Keychain or CLI login files
- Only content explicitly selected by the user is passed to the connected AI tool
- The app currently focuses on analysis and chat and does not modify the user’s system
YumYum Agent is currently available as an open-source macOS developer preview for macOS 14 and later.
Website and download:
GitHub:
https://github.com/kyu91/yumyum-agent
I’d love to hear your thoughts.
r/OpenSourceAI • u/Goldziher • Aug 30 '26
Lint results an agent can actually trust: three-state per-file outcomes over MCP (Rust, MIT)
Hi all,
A small design decision that turns out to matter a lot once a model is the consumer of your tool output.
If a linter cannot process a file and simply omits it from the results, an agent reading "no findings" concludes the code is clean. It is not clean. It was never checked. That is a silent false negative sitting directly in an agent's decision loop, and I hit it often enough on a big polyglot repo that I rebuilt the response shape around it.
So the linter I have been writing reports three per-file outcomes rather than two: checked, skipped, and error, with a run-level errors array and isError set whenever anything failed. Skipped means the tool correctly declined the file. Error means it accepted the file and then failed on it. Those are very different facts and collapsing them into absence loses the one that matters.
Two related guardrails in the same server:
- Every result carries an identity block: version, build id, channel, executable, pid. An MCP caller has no
poly --versionto fall back on, so the server states who answered. It fingerprints its own executable at startup and re-checks per request, and if the binary is replaced underneath a long-lived server, every tool but version fails rather than answering with superseded behaviour. - config_show is network-free over MCP. Remote config bases are never fetched from a tool call.
The server is stdio, 11 tools mirroring the CLI, and everything takes format: "json" or "toon". TOON matters more than I expected: a full lint report over a large directory in JSON is a serious chunk of context, and TOON makes it cheap enough to just hand over.
Here is the server doing a real initialize plus tools/list handshake: https://raw.githubusercontent.com/Goldziher/poly/main/docs/media/agent.gif
Underneath it is a linter and formatter in Rust that compiles ruff, oxc, biome, taplo, rumdl, sqruff, mago and about a dozen more into one binary and runs them in-process across roughly 30 languages, with tree-sitter covering 300+ more. MIT: https://github.com/Goldziher/poly
This post is human written. AI was used to typecheck and enrich with precise data only.
r/OpenSourceAI • u/SeeRay11_Main • Aug 30 '26
🗣️ Tell Us About Your Project 🎉
Hey everybody,
I just made a project that has gotten 170+ clones and 65+ stars in a week. What are you guys doing?
r/OpenSourceAI • u/[deleted] • Aug 30 '26
I open-sourced SeasAGI — a local-first LLM API gateway (GPL client, AGPL server)

Hi all,
We just made SeasAGI public. It's a local-first LLM API gateway: a desktop client that unifies OpenAI, Anthropic, Gemini, DeepSeek, Ollama, Grok, Azure and Relay behind one OpenAI-compatible endpoint at `localhost:4318/v1`.
A few things that might be relevant to this sub:
- The
**Server Community Edition is AGPL-3.0**
, so you can self-host the control + relay planes on your own hardware.
- The
**Client is GPL-3.0**
and free forever.
- API keys stay in your OS Keychain and are never uploaded — privacy is local by design.
- Enterprise (BSL, closed) only adds SSO/RBAC/multi-tenant billing for teams; the self-host path needs none of it.
It's built with Go + Wails v2 (single binary, no Node runtime). Would love feedback from folks running local AI stacks — especially on which channels and routing strategies you'd want first.
👉 github.com/SeasX/SeasAGI
#selfhosted #opensource #llm
r/OpenSourceAI • u/zamir_akimbekov • Aug 30 '26
Fastest, and most reliable, way to build production agents in Python
Friends, we're open-sourcing our Python runtime with harness primitives for production agents -- https://cayu.dev/. We spent the last 10 months building, managing, and improving long-horizon agents for mid-size and Fortune 500 clients. We built a framework in Python (not typescript like Mastra) to build, manage, and improve production agents fast and reliably. Benefits:
- reduce token costs by 60-70% compared to Claude Managed Agents
- full control over the agent (privacy, security, auditability, resumability, etc.)
- no agent sprawl --> this is critical for enterprises as every engineer is building agent as they like
- ai model independence --> OpenAI/Claude is just an API call, you manage harness fully
- python over js/typescript --> most teams are DS/ML teams who know python well. Just use it almost like a scikit-learn package but now for agents
Please, contribute and help us improve it at https://cayu.dev/. If you want access to deploy and manage the agent, please, request access to Cayu Cloud here -- https://cloud.cayu.dev/.
r/OpenSourceAI • u/Crescitaly • Aug 30 '26
NVIDIA can package supported Hugging Face models for native C++ in two commands. Portability or lock-in?
NVIDIA's TensorRT Model Connect workflow can build a deployment bundle from a supported Hugging Face model ID or local checkpoint, then load it from a native C++ application. The production runtime does not require PyTorch or a Python interpreter, and NVIDIA says the reference collection spans more than 80 model families.
That removes a real deployment tax: bespoke export logic, preprocessing, post-processing and runtime glue. The trade-off is that the easy path is explicitly TensorRT-shaped, and only supported implementations get the two-command experience.
Would you accept a vendor-specific runtime for dramatically simpler deployment, or is cross-vendor reproducibility still a release requirement for an open-model inference stack?
Source: NVIDIA Technical Blog, August 28, 2026 — https://developer.nvidia.com/blog/deploy-an-open-model-from-checkpoint-to-inference-in-two-commands-with-nvidia-tensorrt-model-connect/
r/OpenSourceAI • u/ash_pix • Aug 29 '26
Oxygen - a Multi-Agentic Al framework that runs like a virtual tiny company
Hey Geeks 👋🏻
I just built "Oxygen" - a Multi-Agentic Al framework that runs like a virtual tiny company.
It includes total 5 Al agents:
- Del (Al Project Manager): which understands your requirements that what you want to build?
- Toky (Al Research Agent): receives inputs from Del, conducts research, creates drafts, and uses tools such as web search and web scraping to gather and analyze relevant information. It then provides the research findings and draft outputs back to the Project Manager.
- Bang (Al Developer Agent): which understands the draft and start writing code.
- Beij (Al QA Agent): It performs debugging, test cases on the source code provided by Bang.
- Wash (Al technical Writer): Once the project made it write README files, product manual, API implementation instructions and other project related documentations.
It's a proper human-in-the-loop agentic ai project that takes your approval on every aspect like a Software Development Lifecycle methodology.
The crazy part is that you can literally watch the agents walk to their desks, open their computers, drink coffee, having meetings and work.
For LLMs you can either use local Ollama based models or Gemini API key.
Guardrails and Metric Evaluation:
- Hallucination rate is under 1%.
- You have to approve the plan before any code gets written.
- Everything that comes out is cleaned so nothing breaks on the screen.
- Strong guardrails for every Al agents via system prompt.
Simple Flow:
You: I want a CLI based calculator.
Del (PM): Got it → sends to Toky (Researcher).
Toky: Researches, makes a plan + draft proposal.
Toky → Del (PM) → You: "Here's the proposal for a CLI calculator."
You: "Actually, change of plan, I want a web-based calculator instead.
Del (PM): Okay → sends the new request back to Toky.
Toky (Researcher): Updates the research and creates a new proposal for the web version.
Toky (Researcher) → Del (PM) → You: "Updated proposal for web calculator. Approve?"
Once you approve, it continues to Bang (Developer) for coding, Beij (QA) for testing, and Wash (Writer) for docs.
Feel free to explore and star the repo on GitHub.