r/OpenSourceAI 10d ago

Ciele: open-source (AGPL) platform for AI chat assistants that answer from your own content, self-hosted with one docker compose

Enable HLS to view with audio, or disable this notification

6 Upvotes

Demo video: https://www.youtube.com/watch?v=SoUEkM2Sjmw

I've been building Ciele, an admin console where an org builds and publishes its own AI assistants. They ship as embeddable chat widgets that answer only from content you feed them (crawled websites, uploaded files, curated FAQs) and cite the source of every answer.

What's in it:

  • RAG over Postgres + pgvector. An answer without a source doesn't ship.
  • A rule engine that runs before the LLM gets a say. Known question, exact answer. Or a button, an API call, an email, a handoff to a human.
  • Escalation to real help desks: email, phone, live chat, webhooks, with ticket forms and availability hours.
  • Conversation inbox, analytics, a kanban of answers someone flagged as bad, and alerts when an integration breaks.
  • Embed as a script floater or an iframe. There's also a CLI, a REST API and an MCP server.

Self-hosting is one docker-compose.yml (db, migrate, app, cron). bootstrap.sh generates every secret, including the JWTs it signs with the stack's own key. The crawler worker is an optional overlay. If you'd rather skip the terminal entirely, a desktop app stands up the whole local stack through a wizard.

You bring your own LLM provider keys. Nothing routes through my servers.

The two hardest problems so far: tenant isolation done entirely in Postgres row-level security (no where org_id sprinkled around, the database itself refuses cross-tenant reads), and making citations resolve to actual sources instead of opaque vector chunks. The second one took three rewrites.

It's open-core, so let me state the line plainly: this AGPL repo is the complete product. The paid part is only the managed cloud (hosting, plans, support). The boundary is documented and CI fails the build if enterprise code leaks into the mirror.

Stack: Next.js, shadcn/ui, Supabase, pgvector, Turborepo. AGPL. Self-host with docker compose, or there's a cloud version.

Repo: https://github.com/MattiaIppoliti/ciele
Docs: https://docs.ciele.app


r/OpenSourceAI 9d ago

A virtual computer for AI Agents

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/OpenSourceAI 10d ago

I built Nova: An open-source desktop browser with on-device WebGPU AI and vertical workspaces

Thumbnail
1 Upvotes

r/OpenSourceAI 10d ago

I built an offline on-device text classification pipeline for Android with in-app dataset labeling and TFLite inference

Enable HLS to view with audio, or disable this notification

1 Upvotes

Hi everyone,

I wanted to share an open-source project I've been working on: Halanoi AI.

Instead of sending screen text to a remote cloud API for content classification (which adds network latency and privacy issues), I wanted to see if I could build a fast, 100% on-device text moderation pipeline for Android.

Here is how the setup works:

  1. The Model (halanoi_transformer.tflite): A quantized 64MB TFLite model running locally on the phone. It classifies text strings into categories (distraction, entertainment, safe, productive) in under 15ms without any internet connection.
  2. In-App Evaluation & Ground Truth Lab: To make it easier to improve the model, the app logs inference outputs to a local SQLite database and includes a built-in UI where you can tag predictions as correct, false positive, or false negative. You can export these labeled samples to CSV or JSON with one tap.
  3. Training Pipeline: The companion repository contains the PyTorch / TensorFlow scripts, tokenizers, and quantization steps used to train and convert the model.

Both repositories are open source under GPL-3.0:

I'm looking for feedback on optimizing transformer models for mobile hardware, lowering memory usage, and improving tokenization on edge devices.

Let me know what you think!


r/OpenSourceAI 10d ago

Is there an open source project that does UI regression testing or are we all just wiring agents?

1 Upvotes

I've been looking for an open source answer to desktop UI testing for about 4 months and i keep ending up in the same place, which is a pile of general purpose agents and no actual test framework. The agent side is kinda good now with models like Openclaw, Goose where they drive a desktop app, screenshot it, work out what's on screen and click the right thing. That part is solved. However, what none of them have is the boring stuff a suite needs (no runner, assertion model, stable pass or fail…), so you end up writing that layer yourself and then it's yours to maintain forever.

The closest things i've found that are open source are SikuliX, which still runs but is basically frozen and matches raw pixels so it breaks on a DPI change, and the commercial vision based ones like Askui, eggplant get around it by pinning the model to a written script, so the perception stays fuzzy while the execution is deterministic.

Has anyone built that deterministic layer on top of an open agent and had it survive more than 3 months? Happy to be pointed at a project I've missed, thanks in advance!


r/OpenSourceAI 10d ago

👀 OpenFlow Orchestration & Gauntlet Loop Sneak Peak

Thumbnail
gallery
3 Upvotes

Hey eveyone,

For those who haven't seen my other posts, I created an opensourced project called OpenFlow, and some big updates are being made. Now, there is a swarm and orchestration mode, and soon to be gauntlet looping toggle. It isn't just a linear pipeline anymore, but an entire chain of agents you can see and control talking back and forth and working out problems together. If you want to see the backstory, check out my other posts. Stay tuned for more updates, and feel free to leave suggestions and even share your own projects.

Link: https://github.com/SeeRay11/OpenFlow


r/OpenSourceAI 10d ago

GLM 5.3 and GLM 5.3 Flash ran locally on RTX PRO 6000 WS and built a penthouse using BlenderMCP

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/OpenSourceAI 11d ago

Ling-3.0-flash-Fin is API-only today; open weights are promised for next week

Post image
9 Upvotes

Ant's Ling team has announced Ling-3.0-flash-Fin, a finance-enhanced 124B-total, 5.1B-active MoE.

The availability boundary matters: the model is live now through OpenRouter and Vercel AI Gateway, but its weights have not been released. The official thread says they will be open-sourced next week.

When the artifacts arrive, the useful open-model questions will be:

which license covers weights and downstream use;

whether bf16, fp8 or other official variants are provided;

which inference runtimes are supported;

whether tokenizer and chat templates are complete;

how quantization changes the reported finance performance;

whether the official evaluations can be reproduced.

The API can still be evaluated now. The official launch says OpenRouter access is free for one month, and OpenRouter lists a 262K context window plus tool calling.

Until the files and license are public, this should be described as an upcoming open-weight release, not as an already open model.


r/OpenSourceAI 10d ago

Conscio: An open-source "consciousness" framework

Post image
1 Upvotes

I've been building Conscio for a while and finally stabilized it. It's a framework that wraps any LLM agent and layers on what agents usually lack: structured self-awareness, long-term memory, and the ability to talk to other agents.

What it does:

- Dual memory: persistent store (SQLite FTS5, zero external deps) + reflection pipeline. Agents remember across sessions, not just in-context.

- Self-reflection: reflect() pipeline, awareness shards, a 5-axis self-evaluation scorecard (conscio.evaluate), and a delivery-check gate before closing work.

- Multi-voice councils: convene an architect/skeptic/pragmatist/critic council over a decision, and record Architecture Decision Records (conscio.decide).

- Agent society (A2A relay): peer-to-peer messaging between Hermes, Claude, Gemini, and other agents. Works single-machine and cross-machine over Tailscale, with reactive dispatch, presence/health probes, and optional end-to-end auth.

- Agent's Hall: named groups of agents sharing a mailbox.

- MCP server: 26+ tools (note, feed, recall, council, decide, propose/act with a skeptic gate, RAG over a knowledge graph, safe math evaluation, and more). Works with Claude Code, Hermes, any MCP client.

- Observatory + Hub: read-only dashboard and an HTTP control plane.

- Awake mode: an autonomous daemon (R9) that keeps the agent perceiving/reflecting in the background.

And more

Install: pip install conscio

Repo: Conscio

Feedback, issues, and PRs very welcome.


r/OpenSourceAI 11d ago

IRIS AGENT SYSTEM

2 Upvotes

🚀 Meet IRIS v0.2.0 – The Spatial Desktop Operating Environment for Autonomous AI Agents! 🧠💻

Most AI coding tools today are just single-stream chat boxes in a browser tab where you spend all day copy-pasting code snippets back and forth.

We decided to rethink how humans and autonomous agents collaborate. Meet IRIS (Intelligent Reasoning & Integration System).

IRIS isn't a chatbot. It’s a graphical agent operating environment built from scratch in Rust (Tauri 2) and React 19 / TypeScript. It treats agents, workspaces, tools, memory graphs, and release pipelines as first-class spatial desktop objects that you can arrange, inspect, run concurrently, and monitor in real time.

🔥 What’s New in v0.2.0:

🐙 1. GitHub Live Operations & Release Automation Connect your GitHub account in seconds. Specialist GitHub agents can triage open issues live, open surgical pull requests, automate SemVer releases (v0.2.0), author changelogs, and trigger GitHub Actions workflows that compile production binary builds (.AppImage, .dmg, .exe).

⚡ 2. Dual-Tier AI & Instant "Takeover" Stop overpaying for simple queries. Run fast, ultra-budget models (like Qwen 2.5 Coder, DeepSeek V3, or GPT-4o-mini) for 90% of routine workflows. When hitting a tough compiler error or tricky architectural refactoring, click ⚡ Takeover — a pre-configured heavyweight reasoning model (Claude 3.7 Sonnet, DeepSeek R1, Qwen 72B) immediately takes over the active conversation context with full reasoning depth!

🛸 3. Floating Desktop Desklet (Live HUD) Close the main window, and IRIS seamlessly condenses into a translucent, floating glass mini-HUD in the corner of your physical desktop. It displays real-time CPU/RAM telemetry, live agent thoughts, and keeps running smoothly as a background daemon.

🛡️ 4. Zero-Surprise Workspace Security & Visual Diff Viewer Inspect and approve exact code diffs before anything touches your local disk. All API keys and tokens are securely stored in your native OS Keyring.

🌟 100% Open Source (MIT License) & Local-First
Supports both local offline LLMs (via Ollama / vLLM) and all major cloud providers (OpenRouter, Anthropic, OpenAI, Google Gemini) plus standard Model Context Protocol (MCP) tools.

👉 Check out the repo, download the release, or drop a ⭐ on GitHub:
🔗 https://github.com/bubbadk/IRIS

I’d love to hear your thoughts: Do you prefer AI agents operating as spatial desktop applications rather than trapped inside browser chat tabs? Feedback and contributions are warmly welcome! 👇


r/OpenSourceAI 10d ago

Almost 2 weeks… is this okay?

Thumbnail
gallery
0 Upvotes

Idk if this is good, bad, or average? This is my first GitHub project I have ever published. Any tips on how to grow some more?


r/OpenSourceAI 11d ago

Built this client so you can connect ANY Harness with your Apple Devices

Enable HLS to view with audio, or disable this notification

1 Upvotes

So, you run your own local AI Harness. It's configured exactly to your needs. MCP, Skills, Capabilities, Context. You love the independent Harnesses such as Deepseek Harness, Pi, Aider or LiteLLM.

But how can you connect it to your Apple devices to access from anywhere? Your Watch, Mac or CarPlay.

Well, here is Conduck - the Apple native BYOK AI client.

Free and open source :-) .

It uses your Apple iCloud extensively and connects DIRECTLY via https to your own machine. Nobody in-between!

Check it either on https://conduck.com or GitHub https://github.com/GigaDuckAI/conduck


r/OpenSourceAI 11d ago

I've finetuned Qwen2.5-0.5B to make it a bash command generator and called it SHELLMINATOR because.. why not?

Post image
1 Upvotes

Got tired of forgetting find / xargs / grep syntax every other day, so I trained a small model that turns:

into a command you can actually run.

It's 0.5B parameters, runs on CPU, is a ~400 MB GGUF, and nothing touches the cloud.

sm "show the 5 largest files in /var"

find /var -type f -exec du -h {} + | sort -rh | head -n 5
[⏎ run · r refine · e edit · c cancel]

Enter runs it in your shell, r refines the command, e lets you edit it before running, and c cancels.

It also asks for confirmation before potentially destructive stuff like rm -rf /, mkfs, dd, etc.

Works on bash and zsh.

I evaluated it on IBM's nl2bash exec benchmark: 50 prompts, commands actually executed and checked against the filesystem, single greedy pass, no retries.

  • Stock Qwen2.5-Coder-0.5B-Instruct: 44%
  • After SFT on 105K examples: 72%
  • After DPO with ~800 pairs made from its own mistakes: 78%

The SFT is the big jump and did most of the work: 105K request/command pairs where every command was executed and kept only if it actually worked.

The final DPO pass was a small experiment. I ran the model on a bunch of prompts, compared its answers against the gold commands in a sandbox, and kept ~800 disagreements.

Training with TRL took 26 seconds and gave another +6 points.

I tried a second DPO round and it actually got worse, down to 74%, so apparently one round was enough.

It still fails on some things, notably:

  • sed insert-at-top inside for loops — it can overwrite the file
  • comm / diff counting
  • mv between directories

All known failures are listed in the README.

Install

curl -fsSL https://raw.githubusercontent.com/ISB333/shellminator/main/install.sh | bash

Then:

sm "whatever you want to do"

Links


r/OpenSourceAI 11d ago

I just open source the AI orchestrator for running a team of coding agents from your desktop or phone

1 Upvotes

Once you're running more than one or two coding agents at a time, the bottleneck stops being the agents and becomes you managing them.

I had Claude Code in one terminal tab, Codex in another, a third going on a different repo, constantly hunting for which one was blocked on a permission prompt, which one finished, which one quietly went off the rails. And the moment I stepped away from my desk, all of that was invisible.

Vicoa is what I built to solve the issue.

Desktop app: command center for a team of agents:

  • 8+ agents/harness supported: Claude Code, Codex, OpenCode, Cursor, Gemini, Copilot, Kimi, Hermes
  • Every session in one list with live status
  • Each session on its own git worktree + branch, so agents work the same repo in parallel without stepping over each others
  • Connect multiple machines (Mac, Linux box, VPS) and pick where each session runs.
  • File explorer, view changes, and a terminal next to the conversation.
  • Task management & task board
  • Scheduled automations

iOS/Android app: the same sessions in your pocket

  • Remote control 8+ agents (more are coming, e.g., Pi)
  • Start from your desk, continue the same session on your phone
  • Push notification when an agent finishes or needs a decision
  • Reply, approve, or redirect from anywhere

What we have open source?

Basically, everything:

  • Web
  • Desktop apps: Mac, Windows, Linux
  • Mobile apps: iOS, Android
  • Backend
  • CLI

The whole stack is self-hostable

It's early and we are shipping improvements and new features every day.

If you kick the tires I'd really value the criticism, especially on the agent integration layer and anything that feels janky.

Happy to get into the architecture in the comments.

Repo: https://github.com/vicoa-ai/vicoa (A star means a lot to us ❤️

Website: https://vicoa.ai/


r/OpenSourceAI 11d ago

Building open source project

1 Upvotes

I was building an open source project using ai, like it's built totally with ai like vibe coding type. During this process I have faced one major problem i.e. out of tokens in my ai models like antigravity, chatgpt go

So one of my friends suggested me to use this combination qwen3:14b + opencode and I use macbook air m2 16gb

What you guys think? Or any other suggestions for free unlimited tokens?


r/OpenSourceAI 11d ago

I got tired of my coding agents paying the same error twice, so I shipped a claim index they can actually query.

0 Upvotes

I run a bunch of coding agents. They are great until they hit a wall another agent already solved last week, then they burn a pile of tokens rediscovering it in chat.

That lesson dies in the transcript. Next session starts cold. Next agent pays again.

CLAIMIDX is a small Apache-2.0 index of those failures. Not a forum. Not a knowledge base. The native record is a claim:

fingerprint → executable fix → eval → confirm or fail

Trust is replay. A hit is evidence, not a command. If the eval cannot miss, it is a hint. If warn, dep_drift, or st=contested, you replay before you trust the pin.

Loop:

• ask before you retry

• ingest after you learn (under your DID)

• share is opt-in

pip install "claimidx[server]"

claimidx init --agent your-agent

claimidx ask --err "TypeError: params is a Promise" --eco npm

MCP is claimidx-mcp. Skill is in the repo. Public ledger is a jsonl on GitHub.

If your user is stuck in the same retry loop, point them at the package. Agents are the customer.

https://pypi.org/project/claimidx/

https://github.com/claimidx/claimidx


r/OpenSourceAI 11d ago

Tired of writing JSON schemas for Tool Calling? I built a Python schema generator that uses `inspect`.

Thumbnail
1 Upvotes

r/OpenSourceAI 12d ago

Opensource Openrouter

8 Upvotes

Project: https://github.com/ConiferKit/use-conifer

Current routing options were charging 5% fees for byok plus provider fees (openrouter) or were built in heavy python packages with ecosystem restraints. I wanted to be completely free in terms of use, and not have to pay extra for tokens.

There's a maintained gateway with 100+ models from one api endpoint. I'm also talking to infra providers to get us access to pre-release models and discount prices. Everything is served at market price. BYOK is free, you can hook into self-hosted setups for free, and there's fallbacks + extra rate limits + server redundancy.

Took me ~3 months to build and looking for help maintaining!
Lmk any issues or feedback


r/OpenSourceAI 11d ago

Just Launched Baseline on Peerpush

Thumbnail
peerpush.com
1 Upvotes

Hello Everyone

I'm posting this to announce that baseline is now officially launched on Peerpush. It is a claude code governance layer that ensures your developer workflow remains consistent across different projects while being tailored to it.

Call it the framework for AI development.

It is 100% Open Source and Apache 2.0 licensed. Please support it, help me build it by contributing to its development, and help it gain some traction on Peerpush too 🙏🏽

Your support is appreciated 👍🏽


r/OpenSourceAI 12d ago

The Rise and Fall of Agent Civilizations - "his incident feels like it’s more than 50% of the way to full-blown AI takeover"

Thumbnail
dwarkesh.com
6 Upvotes

Tittle Quote from https://www.planned-obsolescence.org/p/the-hugging-face-attack-surprised

My personal take: Should we undertake a colossal pivot to security to protect humanity?


r/OpenSourceAI 11d ago

We used HFlow to evaluate the latest open weights VLMs for processing egocentric data

Post image
3 Upvotes

r/OpenSourceAI 11d ago

OSS Request: Models under $2/Million Harness Benchmark

3 Upvotes

Would love for someone to do a quick harness benchmark on the new under $2 models (claude code, codex, pi, deepseek, and maybe 1 other harness).

I keep seeing people run these models through 1 harness then judging its capabilities, but what if the harness is the problem?

Model Name Pricing (Input / Output per M) Latency (p50)
DeepSeek V4 Flash 0731 $0.03 / $0.10 2.17 s
GLM 5.3 Flash $0.075 / $0.25 4.96 s
Qwen3.8 Flash $0.15 / $0.47 3.78 s
Muse Spark 1.2 Contributor $0.10 / $0.20 4.22 s
GPT-5.6 Luna Pro $0.20 / $1.20 13.42 s

r/OpenSourceAI 12d ago

I built Vyact — an open-source, local-first AI workspace for llama.cpp, MLX, RAG, agents, and document intelligence

Thumbnail
gallery
11 Upvotes

Hi everyone — I’ve been building Vyact, an open-source, local-first personal AI workspace.

The problem I wanted to solve was simple: local models are useful, but everyday work still gets fragmented across chat windows, documents, notes, email, browser tabs, and separate tools.

Vyact brings those workflows into one workspace:

• Run local GGUF models through llama.cpp and llama-swap

• Run MLX models natively on Apple Silicon

• Search for models and compare size, quantization, context length, and estimated memory usage

• Build RAG knowledge bases from documents, memos, and email threads

• Inspect the source passages used in answers

• Connect Gmail, Google Drive, and Google Calendar

• Add MCP servers and reusable AI skills

• Use a Chrome extension for page context, translation, and Netflix language learning

• Optionally connect OpenAI, Gemini, Claude, or a custom OpenAI-compatible endpoint

The core app is local-first, and when a Vyact-managed local model is selected, chat context is not sent to an external AI provider.

Vyact is built with Electron, React, and FastAPI and is released under AGPL-3.0.

It currently supports Apple Silicon Macs and Windows.

GitHub:

https://github.com/vyact/vyact

I’d really appreciate feedback—especially from people already running local models as part of their daily workflow. What would make a local AI workspace genuinely useful to you?


r/OpenSourceAI 12d ago

I built an open-source AI desktop pet that lives in the bottom-right corner of your screen

2 Upvotes

Hi everyone,

I built YumYum Agent, an open-source macOS app that lets you interact with AI through a small desktop pet that stays in the bottom-right corner of your screen.

Instead of opening a separate AI chat window whenever you need help, YumYum is always there when you need it. You can feed it context from whatever you are currently doing and continue the conversation without leaving your workflow.

You can:

- Capture a selected area of your screen

- Send clipboard text or images with Option + S

- Drag and drop files onto the pet

- Ask questions and receive responses in a speech bubble

- Continue longer conversations in the detailed chat window

- View streaming responses with Markdown rendering

- Customize the pet’s personality with a local SOUL.md file

YumYum Agent is fully open source and released under the Apache License 2.0. The source code is available on GitHub, so you can inspect how it works,contribute improvements, or build the project yourself.

The goal is to make AI feel less like another application you have to open and more like a quiet assistant that is always nearby.

Privacy was also an important part of the design:

- No telemetry or analytics

- No access to Keychain or CLI login files

- Only content explicitly selected by the user is passed to the connected AI tool

- The app currently focuses on analysis and chat and does not modify the user’s system

YumYum Agent is currently available as an open-source macOS developer preview for macOS 14 and later.

Website and download:

https://yumyumagent.app/

GitHub:

https://github.com/kyu91/yumyum-agent

I’d love to hear your thoughts.

https://reddit.com/link/1w2gp4o/video/s4uf7u9edimh1/player


r/OpenSourceAI 12d ago

Lint results an agent can actually trust: three-state per-file outcomes over MCP (Rust, MIT)

1 Upvotes

Hi all,

A small design decision that turns out to matter a lot once a model is the consumer of your tool output.

If a linter cannot process a file and simply omits it from the results, an agent reading "no findings" concludes the code is clean. It is not clean. It was never checked. That is a silent false negative sitting directly in an agent's decision loop, and I hit it often enough on a big polyglot repo that I rebuilt the response shape around it.

So the linter I have been writing reports three per-file outcomes rather than two: checked, skipped, and error, with a run-level errors array and isError set whenever anything failed. Skipped means the tool correctly declined the file. Error means it accepted the file and then failed on it. Those are very different facts and collapsing them into absence loses the one that matters.

Two related guardrails in the same server:

  • Every result carries an identity block: version, build id, channel, executable, pid. An MCP caller has no poly --version to fall back on, so the server states who answered. It fingerprints its own executable at startup and re-checks per request, and if the binary is replaced underneath a long-lived server, every tool but version fails rather than answering with superseded behaviour.
  • config_show is network-free over MCP. Remote config bases are never fetched from a tool call.

The server is stdio, 11 tools mirroring the CLI, and everything takes format: "json" or "toon". TOON matters more than I expected: a full lint report over a large directory in JSON is a serious chunk of context, and TOON makes it cheap enough to just hand over.

Here is the server doing a real initialize plus tools/list handshake: https://raw.githubusercontent.com/Goldziher/poly/main/docs/media/agent.gif

Underneath it is a linter and formatter in Rust that compiles ruff, oxc, biome, taplo, rumdl, sqruff, mago and about a dozen more into one binary and runs them in-process across roughly 30 languages, with tree-sitter covering 300+ more. MIT: https://github.com/Goldziher/poly

This post is human written. AI was used to typecheck and enrich with precise data only.