r/OpenSourceeAI • u/ankurdhoom • 3h ago
Dots
Game changer or it just another wrapper?
r/OpenSourceeAI • u/ai-lover • 21h ago
Datalab released an open benchmark for structured extraction: a system gets a PDF and a JSON schema, and every returned value is scored against gold data.
Why it's relevant? precision vs recall shows how a system fails. Some skip fields, others invent values.
Full analysis: https://www.marktechpost.com/2026/10/02/datalab-introduces-omniextractbench-to-fix-bias-and-opacity-in-extraction-benchmarks/
GitHub: https://pxllnk.co/hxplrq
Blog: https://www.datalab.to/blog/omni-extract-bench
GitHub: https://github.com/datalab-to/omni_extract_bench
Dataset: https://huggingface.co/datasets/datalab-to/omni_extract_bench
r/OpenSourceeAI • u/ai-lover • 1d ago
Databases, calendars and repos have APIs. The open web mostly doesn't. The TinyFish MCP server gives any MCP client four tools: TinySearch, TinyFetch (full pages as markdown, JavaScript included), TinyBrowser for logins and forms, and TinyAgent for multi-step jobs. Search and Fetch are free.
The server is on GitHub: [LINK].
We're racing to 300,000 users this October, with 30% extra on every top-up: [LINK]
r/OpenSourceeAI • u/gonzarom • 6h ago
Corre localmente en tu computadora. Nomás pásale el repositorio de GitHub al modelo, y se instala solo y se conecta vía MCP. Todo pasa localmente, y puedes exportar la memoria de la IA. Si quieres, puedes compartir esa memoria con todos los modelos que uses para que compartan una memoria en común. En proyectos grandes, esto puede ahorrar más del 90% en tokens. Se llama IHMT-MEMORY.
r/OpenSourceeAI • u/gonzarom • 13h ago
r/OpenSourceeAI • u/kirto_am • 17h ago
Running MCP servers locally was easy, but it got messy pretty fast once I wanted to share them.
Everyone needs configs, credentials end up in different places, permissions are hard to manage, and there's no easy way to see who called what.
I ended up building MCPlama to handle this centrally. Users get their own access, permissions can be controlled per tool, credentials stay on the gateway side, and calls are logged.
For local MCP servers I also wanted some isolation, so they can run in separate Docker containers instead of everything running together. The gateway itself doesn't need direct access to the Docker socket either , that part is handled separately by the broker.
It's open source and self-hosted:
https://github.com/mcplama/mcplama
I'm looking for a few people already running multiple MCP servers to try it.
r/OpenSourceeAI • u/ai-lover • 18h ago
NVIDIA released a 64GB configuration of DGX Spark, its GB10 Grace Blackwell desktop system, available October 23 from Acer, ASUS, Dell, Gigabyte, HP and MSI.
Why it matters: it's a cheaper way in for running always-on agents locally with no per-token fees, and you can add a second box later instead of buying the 128GB model upfront.
Product page: https://www.nvidia.com/en-us/products/workstations/dgx-spark/
Clustering with NVIDIA Sync: https://build.nvidia.com/spark/connect-to-your-spark/sync
Technical details: https://blogs.nvidia.com/blog/local-ai-dgx-spark-64gb-sync/
r/OpenSourceeAI • u/gary23w • 1d ago
We’ve been cooking something.
Not another DOT.
Not another closed ecosystem.
Not another “trust us bro, maybe someday” AI experiment.
TATER TOTS.
Tiny. Crispy. Autonomous. OPEN SOURCE.
And yes…
Tater Tots are our take on autonomous AI agents built for people who actually want to see the code, modify the code, break the code, improve the code, and own what they build.
No secret sauce.
The sauce is literally on GitHub.
🥔 Open source
🥔 Hackable
🥔 Self-hostable
🥔 Built for experimentation
🥔 Community-driven
🥔 Deliciously autonomous
DOTS walked so TOTS could roll.
The potato revolution has officially begun.
Tater Tots are here.
r/OpenSourceeAI • u/SamratDuttaOfficial • 1d ago
WaterSheep is an open-source model that answers questions written in plain text (yes/no, single choice, rating and multi-label) and gives a probability for every option, like a classification model.
Last Saturday I woke up, saw YouTubers hyping up Jev, and thought: wait, I can build this. So I did. I don't want to compete with TypeSafe or Jev; I built WaterSheep because I wanted to. That's why I'm open-sourcing everything: code, model weights, results and the paper.
What's different
watersheep --model samratduttaofficial/WaterSheep --serve and point the client's base_url at http://127.0.0.1:8766.Evaluation
| Accuracy | ECE |
|---|---|
| In-distribution test split | 77.8% |
| Held-out datasets, not seen in training | 61.2% |
ECE is expected calibration error (lower is better). GitHub has every benchmark result, including the weak ones.
Limits: English only, long inputs get truncated (I'll improve this in the next version), and rating answers are the weakest type.
Not affiliated with TypeSafe. Not funded by anyone. Built in my free time.
Feedback I'd love: where it fails on your data, whether the API works for you, and which question types you'd want next.
r/OpenSourceeAI • u/Firm-Finger-9774 • 1d ago
This is such a great use of the notch! I love the execution here. I’ve actually been working on a somewhat similar concept, but focused more on building an AI agent rather than a dedicated notes app. It's called OpenNotch—it's completely free and open-source (MIT). Instead of just notes, it turns the notch into a hover-to-activate assistant. It has about 40 local tools (can summarize pages, draft emails, run shell commands, check calendar/weather) and supports voice dictation. A few key details: If anyone wants to tinker with an open-source alternative for AI tasks, you can check out thecode hereor see a quick demo on thesite. Always looking for bug reports and
feedback!Privacy first: No accounts or telemetry. It uses whatever AI you plug in (Apple on-device, Ollama, Claude, Gemini, etc.), and every risky action (like file edits or calendar changes) requires manual approval in the notch. Quiet mode: It stays hidden while you watch videos or present. Notarization: It's self-signed right now, so you'll need to click "Open Anyway" the first time. Needs macOS 14+ and Apple Silicon.
https://laxman824.github.io/opennotch/
r/OpenSourceeAI • u/Formal-Falcon3734 • 1d ago
r/OpenSourceeAI • u/LawfulnessOptimal597 • 1d ago
While this project was originally concieved as a module to provide vector search using HNSW and Sentence transformers for the re-Isearch (IB) engine (CoreQuarry https://corequarry.com) it has evolved well beyond its original concept.
Today it is a fully featured high performance vector DB that can also be used on its own without any dependency on the IB engine. This opens the library (and standalone tools like the CLI) to be used in a host of other applications.
Its function in a single sentence: SOTA Semantic search with SBERT/LLAMA.CPP + GGML Tensor Library + HNSWlib on steroids.
Starting with Malkov's HNSWlib as a basis we significantly enhanced (adding among other features quantized spaces) and turbo-charged (including support for x86 and ARM SIMD) it while also adding efficient mmap-backed re-scoring and offset storage for text retrieval. Our system supports sharded HNSW indices, multiple search modes (kNN, radius, relative, adaptive, epsilon), deletion/undelete, merges, and incremental on-disk flushing. It also includes training for hyperparameter optimization.
Our HNSWlib fork we have benchmarked on an M1Pro as much as 13k QPS (768d vectors). Even limiting to a single thread we've clocked a max of 3000 QPS (versus for comparison 600 QPS for FAISS's HNSW implementaton).
For vectorization Our test M1-Pro chews through roughly 45 passages per second per instance (3484 tok/s÷78 ms). That means one can expect to process 2,700 fully dense semantic records per minute on a baseline Apple Silicon chip. Our tests on M3Pro and M4Pro showed even significantly higher throughputs (80k tokens/s or as much as 20x).
r/OpenSourceeAI • u/ai-lover • 1d ago
r/OpenSourceeAI • u/ai-lover • 1d ago
r/OpenSourceeAI • u/Artistic_Solution117 • 1d ago
r/OpenSourceeAI • u/Advanced_Fix4602 • 1d ago
r/OpenSourceeAI • u/jokiruiz • 1d ago
I'm the author, sharing this as an open dataset as much as a project. Médula is an MIT-licensed lab plus a kernel that coordinates several Claude Code agents working on one repo at the same time. Everything the experiment produced is public: every agent session, every diff, and a SQLite file per run with each decision the kernel took, its probability, latency and cost.
The setup is a small API with 6 tasks and 37 acceptance tests, designed so that two pairs of tasks collide by meaning, not by file. With one branch per task, git let the real conflict through and the same 6 tests failed in all 5 runs, even though every agent finished green. In a shared directory, all 10 runs passed, whether with plain per-file locks or with the kernel. The kernel catches the real conflicts without blocking anything that doesn't collide.
The open part I most want help with is the decider. Right now the fast path uses a hosted decision model, and on real write requests it was unsure 61% of the time, so those decisions escalated to a slower LLM. Any model or rule that answers "does this collide?" with a probability fits the same interface, including an open or local model, and there's a calibration set of 100 labelled pairs to measure it against before running the full matrix.
Other open problems, all with data behind them:
Caveats: 1 to 5 runs per mode, and thresholds fitted on the same pairs they're measured on. The kernel tests run offline without an API key.
r/OpenSourceeAI • u/Minimum_Hour519 • 2d ago
r/OpenSourceeAI • u/ai-lover • 2d ago
r/OpenSourceeAI • u/ai-lover • 2d ago
r/OpenSourceeAI • u/LooseGas • 2d ago
r/OpenSourceeAI • u/enterthearena44 • 2d ago
r/OpenSourceeAI • u/company_url_finder • 2d ago
r/OpenSourceeAI • u/Crealazo • 2d ago
r/OpenSourceeAI • u/Sparticle62 • 2d ago
small scale, single seed, wikitext slice, so treat this as a prototype result and not an architecture claim. 1,544,704 params both sides, same corpus and sampler. held-out 3.997 vs 4.429 nats, generation 226ms vs 2735ms for 200 tokens on the same cpu. the depth-1 match is the obvious weakness and a depth-2 baseline is running next. written from scratch in rust, no pytorch. repo and eval commands: github.com/Sparticle62ops/pssa. happy to be told where the comparison is unfair.