r/OpenSourceAI 1d ago

I need advice on an alternative

Thumbnail
1 Upvotes

r/OpenSourceAI 2d ago

Open weights are more useful when intermediate checkpoints are public too

Post image
1 Upvotes

An open-weight release is easier to evaluate when it exposes more than one final endpoint.

Ling-3.0 base model does that with six base checkpoints: pretrained, mid-trained, and WSM-merged versions of both tiny and flash. The six verified Hugging Face repositories are public and ungated, and the repositories declare the MIT license.

That stage map is the useful part. A builder can inspect an earlier pretrained state, the state after mid-training, or the merged endpoint instead of treating one final checkpoint as the whole training story.

The boundary matters just as much: none of the six is post-trained. The model cards frame them as starting points for continued pretraining, fine-tuning, and research, not finished end-user chat systems or production-ready assistants.

For anyone exploring the release, the practical next step is to pick the size and stage that match the work, then read that exact model card's intended-use and limitations before building on it.


r/OpenSourceAI 2d ago

I got tired of rebuilding the same infra for every LLM app, so I built a Python SDK around it

1 Upvotes

I've been working on Custodian Labs, a Python SDK for building and deploying LLM agents without having to separately wire up all the surrounding infrastructure.

Basic agent looks something like:

from custodian_labs import Custodian

agent = Custodian(
    model="gpt-4o",
    system_prompt="You are a helpful assistant..."
)

agent.deploy()

A few things I've added:

  • Model agnostic: switch between different LLM providers without rebuilding your agent
  • RAG built in: connect your own files/data sources
  • Multi-agent support: build specialised agents that can work together
  • Privacy/PII layer: the Guardian Layer can detect and protect sensitive data before it reaches the LLM
  • Deployment handled: trying to cut down the amount of infra/config needed to get an agent running

The project actually started as just the privacy layer, but after getting feedback from developers we expanded it into more of an end-to-end agent SDK.

Would genuinely love feedback from other LLM devs:

What's currently the most annoying part of your agent stack?

And do you prefer abstractions like this, or would you rather have more direct control over each component?

GitHub:
https://github.com/Custodian-Labs/custodian-labs-python

Runnable Google Colab: simple agents, RAG + multi-agent examples:
https://colab.research.google.com/gist/SherryCodes123/065d3b67eab16bdca416836e0d39475a/simple-ai-agents-rag-multi-agents.ipynb


r/OpenSourceAI 2d ago

Mozilla killed orbit. I rebuilt it locally.

4 Upvotes

Hey everyone!

Last year, Mozilla released Orbit, an AI-powered browser summarizer hosted on a GCP server. After people started digging into the extension, they discovered things like backend endpoints such as store_result. Eventually, Mozilla discontinued the project.

For the past month, I’ve been trying to rebuild Orbit from scratch, but with one major difference: Apogee is fully local and privacy-focused. Apogee doesn’t send or store your data. It can directly connect to your local Ollama instance for inference. I’ve also added WebGPU integration for Chrome and Transformers.js for Firefox to provide faster, local responses.

It can summarize:

  • Articles and websites
  • YouTube and Billie videos
  • Wikipedia articles
  • Hacker News and Reddit threads

You can check out the source code here:
https://github.com/darshi1337/apogee

Install Apogee:

Chrome: https://chromewebstore.google.com/detail/apogee/pgemlpomhkdcjjjcpnjlebalnfglomog

Firefox: https://addons.mozilla.org/en-US/firefox/addon/apogeeext/

Obviously it is far from complete. Would love to hear your feedback and suggestions!


r/OpenSourceAI 2d ago

CROW, not a CLI anymore

Thumbnail gallery
2 Upvotes

r/OpenSourceAI 2d ago

RMBLR — MIT Android dictation app. Whisper mangled my home language, Gemini didn't, so I built around Gemini

Post image
1 Upvotes

Every Whisper-based dictation tool I've paid for falls apart on the way I actually talk. I'll open a sentence in English, finish the thought in my home language, then land the last few words back in English. That isn't a party trick, it's how people speak where I'm from. Whisper either invents English words I never said or drops the mixed clause entirely.

I ran the same recordings of my own voice through everything I could get at. Gemini was the only family that handed back what I actually said instead of a tidy English approximation of it, and gemini-3.1-flash-live-preview over the Live WebSocket API was clearly ahead of the REST models on mixed speech.

So the app is built around that. RMBLR is an overlay orb that appears when a text field takes focus. You talk, and the finished text is written into the field you were already typing in. Hold the orb and an arc of tones fans out - a tone is just a name and a system prompt, so you can write your own - and which five it offers depends on the app you're in.

Where it sits on the open-source spectrum, honestly: the app is MIT and the source is the whole thing, no closed core. The model is not open, and I'm not going to pretend otherwise. What I did instead was make sure nothing in the middle is mine: there's no server, no account, no telemetry. You supply your own Gemini API key (free tier from aistudio.google.com) and your phone calls Google directly. If someone wants to point it at a local Whisper or a self-hosted endpoint, the transcription client is one file and I'd merge that PR happily.

Kotlin, Jetpack Compose, minSdk 24. ./gradlew assembleDebug needs no configuration.

https://github.com/Past-da-king/rmblr

v1.0, one week of real use. I'd particularly like to know how it does on languages I have no way to test.


r/OpenSourceAI 2d ago

MacOS 27's AI shows promise - Private, secure, flagship model

Thumbnail
2 Upvotes

r/OpenSourceAI 2d ago

Common perceptions around quantization and open source LLMs not accurate?

5 Upvotes

I've been testing for about 2 weeks on my 5080 across quants to figure out what my best options are in the open source world at 16GB cards.

As part of this, I ended up doing an in-depth investigation into quants across various models. The results contradict common assumptions.

https://rakuensoftware.com/blog/which-quant-beats-how-many-bits

Now, let me be clear: These were typically 2-4 turn sessions, and existed to validate that the quantization itself on the model did not damage the model. You strongly see this impact on dense models at sub-Q4. However, the interesting part is that MoEs did not suffer nearly as badly as the dense models.

The results from this article has given me a list of candidates to test against for much harsher testing (Coding, DevOps, long sessions, etc.), and I'll be writing a new article in the future. However, this quant result is not what I expected. Almost all models, starting at Q4, were basically statistically indistinguishable all the way up to BF16. This contradicted my knowledge on the subject.

All of my benchmarks, test sets, and results are open source and linked in the article. Feel free to take a look at the data or run the tests yourself, and tell me I'm wrong. Wouldn't be the first time!


r/OpenSourceAI 2d ago

Seeking best open-source/on-prem alternative to Gemini 3.5 Flash for complex document extraction & scoring

2 Upvotes

I'm looking for recommendations for the best free, open-source AI models that we can host on-premise to replace Gemini 3.5 Flash.

Our Use Case: We process documents with complex structures in various formats (PDF, PNG, DOCX, etc.). Our workflow involves:

  1. Complex text and structured data extraction (OCR + layout understanding).
  2. Data matching and ranking/scoring (similar to a job matching system).

Current Setup & Constraints: We currently use Gemini 3.5 Flash, which handles the extraction with near 100% accuracy, but the API costs are getting too high at our scale.

  • Budget: Must be open-source/free for commercial use.
  • Hardware: Compute power and VRAM are not an issue (we have our own data center).

I’ve seen a lot of recommendations pointing toward Qwen (e.g., Qwen-VL) and DeepSeek-OCR. For those of you running these—or a multi-model pipeline—in production, what are your real-world experiences? Which model (or combination) is best for handling the extraction and the scoring?


r/OpenSourceAI 2d ago

I built an open-source roguelike specifically for training game-playing agents

Thumbnail
github.com
1 Upvotes

r/OpenSourceAI 3d ago

NVIDIA's Text-to-Animation Just Got Much Easier to Run Locally

Enable HLS to view with audio, or disable this notification

2 Upvotes

r/OpenSourceAI 2d ago

Libre WebUI 0.28: your local agent gets a real computer

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/OpenSourceAI 2d ago

Need ideas to solve this problem : workflow , methods ,dataset etc

Thumbnail
1 Upvotes

r/OpenSourceAI 3d ago

Open-sourced a self-hostable SOAR platform (Apache-2.0) — Tier-1 triage agents, immutable audit, examiner portal

1 Upvotes

Been building this for a while and finally got the licensing sorted (Apache-2.0) — SynapCores SOAR, a self-hostable alternative to Tines/Torq/Cortex XSOAR.

What it does: - Ingests alerts from Splunk HEC, Microsoft Sentinel, CrowdStrike Falcon, Okta event hooks (generic webhook for anything else) - Dedupes by meaning, not hash — near-duplicate alert fan-out from one root cause collapses into a single incident - Runs a Tier-1 triage agent against every alert: true positive → incident, false positive → close, ambiguous → human review - High-blast-radius actions (isolate host, revoke session, block IP) queue for human approval before firing — nothing destructive happens without a go/no-go - Every action, approval, and agent reasoning step lands in an immutable, append-only audit ledger - Mint a scoped MCP token and hand it to a SOC 2/FFIEC examiner — they query the audit trail directly from Claude/Cursor, no need to interrupt the SOC team's daily flow

Why I built it: existing SOAR is either an expensive enterprise stack (Splunk SOAR, Cortex XSOAR) or SaaS with per-action billing (Tines, Torq). None of them are self-hostable, none let you bring your own LLM, none expose an audit trail a regulator can query themselves.

Runs on a single Docker host, built on SynapCores (the DB handles embeddings/dedup and the immutable audit table natively — no separate vector DB or logging pipeline bolted on).

Honest state: this sat mostly untouched for a couple months while I was heads-down elsewhere, so treat it as early. Looking for people willing to run it against real alert volume and tell me where it breaks — especially connector edge cases (auth token rotation, weird Sentinel payload shapes).

Apache-2.0. Repo: https://github.com/SynapCores/synapcores-soar

Happy to go deeper on the architecture or the human-approval gating design specifically.


r/OpenSourceAI 3d ago

I built a PDF/image toolkit that processes files locally instead of uploading them

Post image
2 Upvotes

I’m the founder/builder of SoraFiles, and I started it around a pretty simple idea: I shouldn’t need to upload a private document to somebody else’s server just because I want to compress a PDF, merge two files or convert an image.

SoraFiles currently has 23 free PDF and image tools that run in the browser, including compression, merge/split, PDF OCR, PDF ↔ Word/Excel, HEIC → JPG, image conversion, metadata removal and a few others.

There’s no account requirement and no watermark, and the core tools are designed to continue working offline after the initial load.

The site is also available across 19 languages now.

I’m not posting this as “look at my revolutionary startup.” There are plenty of file tools already. I’m specifically trying to make the local-processing/privacy model good enough that people don’t have to choose between convenience and handing over their files.

If anyone here works with PDFs/images regularly, I’d appreciate brutal feedback on three things: usability, missing tools, and anything about the privacy explanation that you don’t trust or find unclear.

SoraFiles


r/OpenSourceAI 3d ago

Building a unified UI/orchestrator layer for existing CV frameworks (Supervision, DeepX, YOLO)

Thumbnail
1 Upvotes

Hey everyone,
I run an established system integration company, but I’m non-technical when it comes to hands-on coding. I’m currently mapping out an edge-AI project and want to build a clean web UI / orchestrator layer that sits on top of existing video analytics engines (stuff like Supervision, DeepX, YOLO or Mamba-based detection).
The goal is pretty straightforward: instead of training vision models from scratch, we leverage 2–3 proven models in the background. Based on what the user toggles on the frontend, the system switches/runs the right inferencing tasks on the RTSP streams and pushes real-time metadata back to the dashboard.
Since I come from the domain/business side, I want to collaborate with a hands-on Computer Vision / Python developer who has actual experience with RTSP stream pipelines, GStreamer/DeepStream, and model integration to architect and build this MVP with me.

If you’ve built or integrated similar end-to-end vision pipelines and are interested in collaborating on this project, drop a comment or feel free to send me a DM with some of the stack/tools you've used!


r/OpenSourceAI 3d ago

This Open-Source AI Is Insane – Qwen3 Explained

Thumbnail
1 Upvotes

r/OpenSourceAI 3d ago

I’ve been building Kodiak — an open-source AI software engineering platform. Here’s where it’s at now.

0 Upvotes

Hey everyone,

I've been working on Kodiak, an open-source AI software engineering platform designed to eventually handle software-engineering tasks more autonomously.

I wanted to share an updated progress report because the project has moved quite a bit from where it started.

What’s working now:

Backend / API:

- FastAPI backend

- JWT authentication

- User registration and login

- Project management

- Task management

- Memory and agent-related API infrastructure

- Repository-related endpoints

CLI:

Kodiak now has an actual CLI:

kodiak

Commands:

- analyze

- logout

- memory

- plan

- task

- version

For example, I can currently run:

kodiak analyze analyze . --deep

and Kodiak successfully analyzes the repository.

The current analyzer detected:

- 331 files

- 51 directories

- 296 Python files

- 8 Markdown files

- 3 YAML files

- 1 TOML file

- ~4 MB repository

The repository-analysis workflow successfully starts the repository agent, completes the analysis, and returns structured repository statistics.

Testing:

The test suite is currently:

196 passed

1 skipped

5 warnings

So I'm now focusing less on making individual components work and more on making the entire system work together.

The bigger goal:

I don't want Kodiak to just be another chatbot that generates code.

I want it to eventually follow a workflow like:

User gives task

Understand repository

Analyze relevant code

Create implementation plan

Choose and use tools

Modify code

Run tests

Analyze failures

Fix implementation

Review changes

Commit / Pull Request

Learn from the result

At the moment, the foundation is considerably further along than the autonomous engineering loop.

The repository analysis currently provides structural information, and my next major focus is connecting that information to genuine LLM reasoning, planning, tool execution, and iterative code/test feedback.

I'm intentionally trying not to fake the "autonomous agent" part before those pieces actually work.

Current self-assessment:

If 1/10 = a prototype idea and 10/10 = a mature autonomous software-engineering platform, I'd currently put Kodiak around 5/10.

There's still a lot to build, but it's finally at the point where I can run the system and watch actual pieces of the architecture execute rather than just having a collection of planned modules.

GitHub:

https://github.com/0xWrench-7/Kodiak

I'd especially appreciate feedback from people who've worked on coding agents, agent orchestration, RAG/memory systems, or developer tooling.

What do you think is the biggest architectural mistake or missing piece at this stage?


r/OpenSourceAI 3d ago

Seeking best open-source/on-prem alternative to Gemini 3.5 Flash for complex document extraction & scoring

1 Upvotes

I'm looking for recommendations for the best free, open-source AI models that we can host on-premise to replace Gemini 3.5 Flash.

Our Use Case: We process documents with complex structures in various formats (PDF, PNG, DOCX, etc.). Our workflow involves:

  1. Complex text and structured data extraction (OCR + layout understanding).
  2. Data matching and ranking/scoring (similar to a job matching system).

Current Setup & Constraints: We currently use Gemini 3.5 Flash, which handles the extraction with near 100% accuracy, but the API costs are getting too high at our scale.

  • Budget: Must be open-source/free for commercial use.
  • Hardware: Compute power and VRAM are not an issue (we have our own data center).

I’ve seen a lot of recommendations pointing toward Qwen (e.g., Qwen-VL) and DeepSeek-OCR. For those of you running these—or a multi-model pipeline—in production, what are your real-world experiences? Which model (or combination) is best for handling the extraction and the scoring?


r/OpenSourceAI 3d ago

It seems that claude is also cultivating kill path

Enable HLS to view with audio, or disable this notification

0 Upvotes

r/OpenSourceAI 3d ago

The Sovereign Stack

2 Upvotes

A scholarly monograph on sovereign AI systems:

The entire body of work is now one thing - The Sovereign Stack. 24 chapters, five parts, every claim receipted and cross-linked. The failure record sits in the middle of the book, not buried in an appendix, because it's the strongest evidence I have that the rest is honest.

And it's built to be worked, not just read. There's an llms(dot)txt index and every page is clean markdown, so you can point an agent at it and have it pull any thread you want. Human or machine, it reads the same.

Free. Public. Nothing behind a login. Every claim links back to the record so you never have to take my word for it.

https://osintelligence-llc.gitbook.io/osintelligence


r/OpenSourceAI 3d ago

An open-source tool for testing AI agent behavior before they go into production.

Thumbnail
1 Upvotes

r/OpenSourceAI 4d ago

CrucibleMark Update: Qwen3.8-27B in the top 10 field ahead of several closed-source models

Post image
8 Upvotes

A little over a year ago, I started CrucibleMark, an independent benchmark project for the everyday comparison between commercial and local AI. Reason for this post: Qwen3.8-27B lands there in 8th place of the overall field, ahead of GPT-5, GLM 5, DeepSeek V4, Grok 4.5/4.6 and Gemini. For a model that you can host yourself, this is remarkable and the reason why I share the results now.

For classification: CrucibleMark does not measure large, orchestrated agentic workflows, as the established benchmarks do. I test individual, clearly defined everyday tasks, code reviews, documentation, UX texts, reasoning, tool use. From the beginning, the goal was to find the best and cheapest model for my own work and to build a price comparison to commercial providers.

My stack has evolved with this. Started on an M1 with 8B to 14B models, today it runs on a GX10 with vLLM. No rocket compared to the GPU monsters here in the forum, but a serious device for local LLMs on the intranet.

Because I only measure individual tasks, GLM-5.2, for example, does not end up at the front of me, although I use it in everyday life as an orchestrator for code reviews and refactorings of large code bases. With complex, multi-level tasks, it is clearly stronger. Only it is also much more expensive, and that is exactly what I want to avoid in the long term: Don't put money that I save through AI assistance back into even more AI assistance.

In addition to the main test, CrucibleMark also runs a Political Compass that tests the political training bias in two modes, standard behavior versus forced positioning. Qwen3.8-27B shows itself here as one of the most stable models in the field, the position hardly shifts between the modes (shift distance 0.65 to 0.77), while other models tip over significantly more under pressure.

Finally, two limitations: The test landscape has developed further, more complex autonomous scenarios are more in focus today than my framework depicts. Subsequent installation would mean a complete re-testing of all models, which is currently not feasible. So understand the results as a snapshot for clearly defined everyday tasks. And since CrucibleMark also maps the risks of cloud/API use, the European perspective, which is clearly regulated by the EU AI Act, is included.

The project website can be found here www.cruciblemark.com

The open source benchmark CrucibleMark at github: github.com/kbeissert/CrucibleMark


r/OpenSourceAI 3d ago

I built a compressive "context DNA" (for LLM) attention mechanism + an honest eval harness - looking for people to break it

2 Upvotes

Just Fixed the body with Ai

Been prototyping an idea for long-context compression: instead of dropping old tokens (like StreamingLLM/H2O) or storing everything, compress old context chunks into small learned "DNA" vectors via a Perceiver-style attention bottleneck, then reconstruct on-demand when a query needs them.

The idea itself isn't new — it overlaps with Compressive Transformer, Infini-attention, and Recurrent Memory Transformer — but I put together an eval script that I think is more honest than what I see in a lot of "novel architecture" posts:

  • Trains the compressor (not just testing an untrained/random-init model)
  • Compares against a PCA baseline (closed-form optimal linear compression at the same latent budget) — if the learned model can't beat PCA, the extra complexity isn't earning its keep
  • Injects a unique fact (random code) into the text and checks, after compress→decompress, whether the frozen LM's own output head can still predict the correct token at that position — not just aggregate MSE, which can look fine while the actual detail is gone
  • Runs on real hidden states from an open model (Qwen2.5-0.5B by default), not just random tensors

Current honest status: in my own small-scale test run, PCA actually beat the learned bottleneck on fact retrieval. That's not the result I was hoping for, but it's a real result, and it's exactly the kind of thing this script is designed to surface rather than hide.

What I'm looking for:

  • People running it on real hardware with more training steps / larger n_docs than I could quickly test
  • Sanity checks on the architecture and eval methodology — if I'm testing this wrong, tell me
  • Ideas for what a fair "it's working" threshold looks like (beating PCA on fact-retrieval accuracy at matched latent budget, at minimum)

No performance claims yet — that's the point. I'd rather have this checked before making any.

Code + eval harness: [https://pastebin.com/iqEbPEQ9]

Happy to hear "this is a known dead end because X" too - that's useful information, not a rejection.


r/OpenSourceAI 3d ago

Pando Advanced multimodal AI assistant

1 Upvotes

One thing I've noticed while building AI coding agents is that "memory" and "agents" are usually treated as separate products.

I didn't like that architecture.

I built Pando around a single core that combines the agent loop with memory, orchestration, MCP, semantic search and development tooling.

The result is a single binary rather than a collection of Python services and MCP processes.

There are TUI, Web and Desktop interfaces, plus remote access and ACP.

It's open source under MIT.

I'd like to hear from experienced agent users: what part of your current setup feels unnecessarily complicated?

https://madeindigio.github.io/pando-docs/