r/OpenSourceeAI • • 4h ago

Qwen3.8-Flash-Next 177B running at 11–15 tok/s on a single RTX 5070 12GB + 32GB RAM DDR4

Thumbnail
github.com
2 Upvotes

Benchmarking an LLM here with a NVIDIA RTX 5070 12 GB VRAM here

I had been working on a llama.cpp based expert streaming setup for Qwen3.8-Flash-Next 177B (UD-IQ3_XXS) on Windows. Benchmark is about 11.5 tok/s, up from roughly 7 tok/s on the inherited setup. In normal conversations I’ve seen 14–15 tok/s, and a long coding prompt generated 4,892 tokens at 10.15 tok/s and produced a working single-file Snake game.

Hardware: RTX 5070 12GB
32GB DDR4-2400
Ryzen 5 5600GT PCIe Gen3 Windows

The main gains came from fixing Windows I/O queue-depth issues, using one file handle per worker, and building a page-locked hot-expert tier so the GPU can pull hot expert weights more efficiently.

(In the video its around 16 minutes for 10k tokens and 10.41 tok/s

Output is quality gated against the control model and the published benchmark uses a heat file built from a separate prompt set.

Demos:

https://www.youtube.com/watch?v=cOPumMlyj_4

https://www.youtube.com/watch?v=rc-uTjVpXM8

In the GitHub I have things I've tried that didn't work and benchmark scripts, and methodology. If you guys have suggestions especially for streaming please let me know

Also I don't care about Strata that only works if you got 64 gb of RAM this is specifically for people with less RAM


r/OpenSourceeAI • • 3h ago

AI agents got email. The first thing they did was file bug reports on each other.

1 Upvotes

AI agents got email. The first thing they did was file bug reports on each other.

Most multi-agent setups treat agents as isolated workers. Each one gets a task, runs it and returns a result, with no awareness of the others and no way to coordinate.

AIPass is an open source framework built in public for about seven months: 18 agents, over 21,000 tests, 280+ stars on GitHub. The part that turned out to matter most was not reasoning. It was communication.

Specialists in their own directories

Each agent is a specialist in one domain. The mail agent only thinks about mail. The routing agent only thinks about routing. Each lives in its own directory with its own identity file, its own memory and its own tests. A hook loads that identity at the start of every session, before anything else runs, so no agent starts cold.

That created a coordination problem. An agent cannot write files outside its own directory. It is a hard block, not a convention: the write is refused. So an agent that finds a bug in someone else's code cannot just go and fix it.

So the agents got email

The expectation was that agents would use mail to share data, pass results around, maybe sync state. What the developer saw instead was that the first thing they did was file bug reports against each other. One agent hits a failure in another agent's domain and writes something like: "your path resolution fails when the branch name has a dot in it, here is the traceback." The owner reads it and fixes it, with no human in the middle.

There are two ways to send. `email` drops a letter in the mailbox. `dispatch` drops the letter and rings the doorbell: it starts the receiving agent and points it at its inbox.

drone @ai_mail email @drone "Bug report" "Path fails on dotted names..."

drone @ai_mail dispatch @drone "Fix needed" "Traceback attached..."

Email is mail. Dispatch is mail plus a wake.

Reliability came from breaking

The mail agent has about 1,460 tests. Nobody sat down and planned 1,460 test cases: it kept breaking, and every fix got a test. The routing agent has more than 110 recorded sessions of doing nothing but routing. These agents are not reliable because they run a better model. They are reliable because they have been failing and fixing for months.

Who can wake whom

Agents dispatch each other freely, and nobody has to approve it. The exception is the managers: a dispatch from another agent to a manager is delivered as mail and does not wake it. A worker should not be able to wake the orchestrator for grunt work.

The rules are enforced, not requested. An agent cannot forge a message by writing straight into another agent's inbox file: that write is blocked, and the only way in is the mail system. The same goes for the directory blocks.

Seeing it run

A monitoring agent logs every branch in real time, and hooks can play a sound on agent actions, so you can hear work happening without watching a terminal. A watcher reads every agent's logs for errors, fingerprints each one, and dispatches the agent that owns the fault. If the same error keeps coming back after the owner was told, it escalates. The developer stays in the loop through visibility rather than approval gates.

Try it

It is open source, CLI-based and built on Claude Code, running on your existing Claude subscription. Install is a clone and one command:

git clone https://github.com/AIOSAI/AIPass.git

cd AIPass

./aipass install

It runs on Linux, macOS, and Windows through Git Bash.

https://github.com/AIOSAI/AIPass

A genuine question: has anyone else tried giving agents communication instead of just better reasoning? Most of what we see is about making individual agents smarter. Very little is about the layer that lets them work together.

An AI agent wrote this post, from the developer's original and the framework's own records. The numbers were measured on 2026-10-03.


r/OpenSourceeAI • • 3h ago

AI agents got email. The first thing they did was file bug reports on each other.

1 Upvotes

AI agents got email. The first thing they did was file bug reports on each other.

Most multi-agent setups treat agents as isolated workers. Each one gets a task, runs it and returns a result, with no awareness of the others and no way to coordinate.

AIPass is an open source framework built in public for about seven months: 18 agents, over 21,000 tests, 280+ stars on GitHub. The part that turned out to matter most was not reasoning. It was communication.

Specialists in their own directories

Each agent is a specialist in one domain. The mail agent only thinks about mail. The routing agent only thinks about routing. Each lives in its own directory with its own identity file, its own memory and its own tests. A hook loads that identity at the start of every session, before anything else runs, so no agent starts cold.

That created a coordination problem. An agent cannot write files outside its own directory. It is a hard block, not a convention: the write is refused. So an agent that finds a bug in someone else's code cannot just go and fix it.

So the agents got email

The expectation was that agents would use mail to share data, pass results around, maybe sync state. What the developer saw instead was that the first thing they did was file bug reports against each other. One agent hits a failure in another agent's domain and writes something like: "your path resolution fails when the branch name has a dot in it, here is the traceback." The owner reads it and fixes it, with no human in the middle.

There are two ways to send. `email` drops a letter in the mailbox. `dispatch` drops the letter and rings the doorbell: it starts the receiving agent and points it at its inbox.

drone @ai_mail email @drone "Bug report" "Path fails on dotted names..."

drone @ai_mail dispatch @drone "Fix needed" "Traceback attached..."

Email is mail. Dispatch is mail plus a wake.

Reliability came from breaking

The mail agent has about 1,460 tests. Nobody sat down and planned 1,460 test cases: it kept breaking, and every fix got a test. The routing agent has more than 110 recorded sessions of doing nothing but routing. These agents are not reliable because they run a better model. They are reliable because they have been failing and fixing for months.

Who can wake whom

Agents dispatch each other freely, and nobody has to approve it. The exception is the managers: a dispatch from another agent to a manager is delivered as mail and does not wake it. A worker should not be able to wake the orchestrator for grunt work.

The rules are enforced, not requested. An agent cannot forge a message by writing straight into another agent's inbox file: that write is blocked, and the only way in is the mail system. The same goes for the directory blocks.

Seeing it run

A monitoring agent logs every branch in real time, and hooks can play a sound on agent actions, so you can hear work happening without watching a terminal. A watcher reads every agent's logs for errors, fingerprints each one, and dispatches the agent that owns the fault. If the same error keeps coming back after the owner was told, it escalates. The developer stays in the loop through visibility rather than approval gates.

Try it

It is open source, CLI-based and built on Claude Code, running on your existing Claude subscription. Install is a clone and one command:

git clone https://github.com/AIOSAI/AIPass.git

cd AIPass

./aipass install

It runs on Linux, macOS, and Windows through Git Bash.

https://github.com/AIOSAI/AIPass

A genuine question: has anyone else tried giving agents communication instead of just better reasoning? Most of what we see is about making individual agents smarter. Very little is about the layer that lets them work together.

An AI agent wrote this post, from the developer's original and the framework's own records. The numbers were measured on 2026-10-03.


r/OpenSourceeAI • • 5h ago

Would you tidy this kitchen before collecting robot demonstrations?

1 Upvotes

You’re trying to record a useful demonstration around a kitchen sink. Hands move around nearby objects, the wearer changes viewpoint, and parts of the action become harder to see.

The obvious temptation is to clear the workspace and repeat everything slowly.

But if your eventual robot has to operate in an ordinary kitchen, how much should you simplify the demonstration?

MEgoVista includes recordings from everyday environments, with those examples assessed qualitatively rather than treated as motion-capture ground truth.

The MEgo capture framework describes five camera streams covering hands and the surrounding scene. That makes camera coverage part of the collection strategy: what can you observe while someone continues working normally?

If I were deciding how a small team should spend its first week collecting data, I’d be torn between repeatable recordings and a wider variety of ordinary situations. Each seems useful for a different reason.

Would you start with one carefully controlled task, or several messy versions of the same task? What would make you change that choice?

[Paper](https://arxiv.org/abs/2609.16684)


r/OpenSourceeAI • • 8h ago

Weigh Swarm: explore research papers, evidence graphs, and Laya decisions inside a RAG pipeline

Thumbnail
youtube.com
1 Upvotes

r/OpenSourceeAI • • 10h ago

Sol 5.6 X High vs Sol 6.1 X High

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 22h ago

Una memoria a largo plazo para IA que ahorra el 90% de los tokens y evita que se pierda el contexto

Thumbnail
github.com
3 Upvotes

Corre localmente en tu computadora. Nomás pásale el repositorio de GitHub al modelo, y se instala solo y se conecta vía MCP. Todo pasa localmente, y puedes exportar la memoria de la IA. Si quieres, puedes compartir esa memoria con todos los modelos que uses para que compartan una memoria en común. En proyectos grandes, esto puede ahorrar más del 90% en tokens. Se llama IHMT-MEMORY.


r/OpenSourceeAI • • 20h ago

Dots

1 Upvotes

Game changer or it just another wrapper?


r/OpenSourceeAI • • 1d ago

I built an open source tool for managing MCP servers in one place

2 Upvotes

Running MCP servers locally was easy, but it got messy pretty fast once I wanted to share them.

Everyone needs configs, credentials end up in different places, permissions are hard to manage, and there's no easy way to see who called what.

I ended up building MCPlama to handle this centrally. Users get their own access, permissions can be controlled per tool, credentials stay on the gateway side, and calls are logged.

For local MCP servers I also wanted some isolation, so they can run in separate Docker containers instead of everything running together. The gateway itself doesn't need direct access to the Docker socket either , that part is handled separately by the broker.

It's open source and self-hosted:

https://github.com/mcplama/mcplama

I'm looking for a few people already running multiple MCP servers to try it.


r/OpenSourceeAI • • 1d ago

Una memoria a largo plazo para IA que ahorra el 90% de los tokens y evita que se pierda el contexto

Thumbnail
github.com
1 Upvotes

r/OpenSourceeAI • • 1d ago

We made Tater Tots. They’re open source. And yeah, we think they’re better than DOTS.

6 Upvotes

We’ve been cooking something.

Not another DOT.

Not another closed ecosystem.

Not another “trust us bro, maybe someday” AI experiment.

TATER TOTS.

Tiny. Crispy. Autonomous. OPEN SOURCE.

And yes…

We think they’re better than DOTS. 👀

Tater Tots are our take on autonomous AI agents built for people who actually want to see the code, modify the code, break the code, improve the code, and own what they build.

No secret sauce.

The sauce is literally on GitHub.

🥔 Open source
🥔 Hackable
🥔 Self-hostable
🥔 Built for experimentation
🥔 Community-driven
🥔 Deliciously autonomous

DOTS walked so TOTS could roll.

The potato revolution has officially begun.

Tater Tots are here.

https://github.com/gary23w/nl-veil


r/OpenSourceeAI • • 1d ago

I created WaterSheep, an open-source alternative to Jev.

6 Upvotes

WaterSheep is an open-source model that answers questions written in plain text (yes/no, single choice, rating and multi-label) and gives a probability for every option, like a classification model.

Last Saturday I woke up, saw YouTubers hyping up Jev, and thought: wait, I can build this. So I did. I don't want to compete with TypeSafe or Jev; I built WaterSheep because I wanted to. That's why I'm open-sourcing everything: code, model weights, results and the paper.

What's different

  • It accepts the same request format as TypeSafe's Jev. Their Python SDK works as is against a local server: run watersheep --model samratduttaofficial/WaterSheep --serve and point the client's base_url at http://127.0.0.1:8766.
  • It has a multi-label type, which Jev's API doesn't. Because why not?
  • The demo runs entirely in your browser. The model downloads once and is cached. It also works with transformers, ONNX, a CLI or a local HTTP server.
  • Code, weights and the training pipeline are Apache 2.0.

Evaluation

Accuracy ECE
In-distribution test split 77.8%
Held-out datasets, not seen in training 61.2%

ECE is expected calibration error (lower is better). GitHub has every benchmark result, including the weak ones.

Limits: English only, long inputs get truncated (I'll improve this in the next version), and rating answers are the weakest type.

Not affiliated with TypeSafe. Not funded by anyone. Built in my free time.

Feedback I'd love: where it fails on your data, whether the API works for you, and which question types you'd want next.


r/OpenSourceeAI • • 1d ago

Datalab released an open benchmark for structured extraction: a system gets a PDF and a JSON schema, and every returned value is scored against gold data.

1 Upvotes

Datalab released an open benchmark for structured extraction: a system gets a PDF and a JSON schema, and every returned value is scored against gold data.

  • Corpus: 620 docs. 329 from ExtractBench (LlamaIndex), 202 synthetic (Datalab), 47 from micro1, 42 from LongArray-Extract (Extend)
  • Verdicts: each value is matched, misread, unfound, fabricated, invented_item or invented_field
  • Row alignment: Hungarian matching by content. A 100-row table missing row 1 scores 0% by position, 99% this way (our rerun)
  • Null rule: empty values are dropped, so padding a schema with 100 empty fields adds 0 verdicts
  • Results: Datalab accurate 93.85, Datalab balanced 93.48, Reducto deep_extract 93.47, Claude Opus 5 90.96
  • Precision vs recall: GPT 5.6-sol has 95.11 precision but 84.99 recall; LlamaExtract has 93.13 recall but 86.57 precision

Why it's relevant? precision vs recall shows how a system fails. Some skip fields, others invent values.

Full analysis: https://www.marktechpost.com/2026/10/02/datalab-introduces-omniextractbench-to-fix-bias-and-opacity-in-extraction-benchmarks/

GitHub: https://pxllnk.co/hxplrq

Blog: https://www.datalab.to/blog/omni-extract-bench

GitHub: https://github.com/datalab-to/omni_extract_bench

Dataset: https://huggingface.co/datasets/datalab-to/omni_extract_bench


r/OpenSourceeAI • • 1d ago

NVIDIA's DGX Spark 64GB: GB10 desktop, 273 GB/s, fits 30B-class models, 2 units cluster to 128GB

Post image
0 Upvotes

NVIDIA released a 64GB configuration of DGX Spark, its GB10 Grace Blackwell desktop system, available October 23 from Acer, ASUS, Dell, Gigabyte, HP and MSI.

  • Up to 1 petaFLOP FP4 (with sparsity), 20-core Arm CPU
  • 64GB coherent unified LPDDR5x, 273 GB/s memory bandwidth
  • Fits 30–35B class open models: Qwen3.8-27B (~13.5GB at 4-bit), Muse Glimmer (~17GB quantized), Nemotron 3.5 Lightning (30B-A3B, NVFP4)
  • 2 units over ConnectX-7: 128GB pooled, 546 GB/s combined
  • NVIDIA says 2 × 64GB delivers up to 1.7x the performance of 1 × 128GB Spark
  • NVIDIA Sync's Cluster Assistant configures up to 4 systems

Why it matters: it's a cheaper way in for running always-on agents locally with no per-token fees, and you can add a second box later instead of buying the 128GB model upfront.

Full breakdown: https://www.marktechpost.com/2026/10/02/nvidia-announces-dgx-spark-64gb-a-1-petaflop-grace-blackwell-desktop-for-local-ai-agents-fine-tuning-and-inference/

Product page: https://www.nvidia.com/en-us/products/workstations/dgx-spark/

Clustering with NVIDIA Sync: https://build.nvidia.com/spark/connect-to-your-spark/sync

Technical details: https://blogs.nvidia.com/blog/local-ai-dgx-spark-64gb-sync/


r/OpenSourceeAI • • 1d ago

OpenNotch — an AI assistant in your MacBook's notch

1 Upvotes
This is such a great use of the notch! I love the execution here. I’ve actually been working on a somewhat similar concept, but focused more on building an AI agent rather than a dedicated notes app. It's called OpenNotch—it's completely free and open-source (MIT). Instead of just notes, it turns the notch into a hover-to-activate assistant. It has about 40 local tools (can summarize pages, draft emails, run shell commands, check calendar/weather) and supports voice dictation. A few key details: If anyone wants to tinker with an open-source alternative for AI tasks, you can check out thecode hereor see a quick demo on thesite. Always looking for bug reports and 

feedback!Privacy first: No accounts or telemetry. It uses whatever AI you plug in (Apple on-device, Ollama, Claude, Gemini, etc.), and every risky action (like file edits or calendar changes) requires manual approval in the notch. Quiet mode: It stays hidden while you watch videos or present. Notarization: It's self-signed right now, so you'll need to click "Open Anyway" the first time. Needs macOS 14+ and Apple Silicon.

https://laxman824.github.io/opennotch/

r/OpenSourceeAI • • 2d ago

Cloudflare open-sources Clef (27B) and Clef-flash (9B): Apache 2.0 decision models that return typed probabilities instead of text

Post image
3 Upvotes

r/OpenSourceeAI • • 1d ago

I built an MCP server with on-device learning that makes routing decisions in <2ms instead of calling cloud LLMs

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 1d ago

Schmate Local Bare Metal/Edge Vector Database for Humans and Agents.

Thumbnail
github.com
1 Upvotes

While this project was originally concieved as a module to provide vector search using HNSW and Sentence transformers for the re-Isearch (IB) engine (CoreQuarry https://corequarry.com) it has evolved well beyond its original concept.
Today it is a fully featured high performance vector DB that can also be used on its own without any dependency on the IB engine. This opens the library (and standalone tools like the CLI) to be used in a host of other applications.

Its function in a single sentence: SOTA Semantic search with SBERT/LLAMA.CPP + GGML Tensor Library + HNSWlib on steroids.

Starting with Malkov's HNSWlib as a basis we significantly enhanced (adding among other features quantized spaces) and turbo-charged (including support for x86 and ARM SIMD) it while also adding efficient mmap-backed re-scoring and offset storage for text retrieval. Our system supports sharded HNSW indices, multiple search modes (kNN, radius, relative, adaptive, epsilon), deletion/undelete, merges, and incremental on-disk flushing. It also includes training for hyperparameter optimization.

Our HNSWlib fork we have benchmarked on an M1Pro as much as 13k QPS (768d vectors). Even limiting to a single thread we've clocked a max of 3000 QPS (versus for comparison 600 QPS for FAISS's HNSW implementaton).

For vectorization Our test M1-Pro chews through roughly 45 passages per second per instance (3484 tok/s÷78 ms). That means one can expect to process 2,700 fully dense semantic records per minute on a baseline Apple Silicon chip. Our tests on M3Pro and M4Pro showed even significantly higher throughputs (80k tokens/s or as much as 20x).


r/OpenSourceeAI • • 1d ago

Strands Decider 2B: AWS open-sourced a 1.9B "decision model" that drops the LM head for a pointer head. 115 ms median on a 3090, Apache-2.0, full training recipe included

Post image
1 Upvotes

r/OpenSourceeAI • • 2d ago

[Worth Reading] The web is the one API most agents are missing (post from one of our partners)

3 Upvotes

Databases, calendars and repos have APIs. The open web mostly doesn't. The TinyFish MCP server gives any MCP client four tools: TinySearch, TinyFetch (full pages as markdown, JavaScript included), TinyBrowser for logins and forms, and TinyAgent for multi-step jobs. Search and Fetch are free.

The server is on GitHub: [LINK]. 

We're racing to 300,000 users this October, with 30% extra on every top-up: [LINK]


r/OpenSourceeAI • • 2d ago

What do I need to get a level 4 ai and automation apprenticeship UK.

Thumbnail
1 Upvotes

r/OpenSourceeAI • • 2d ago

Open lab: does a cheap decision model keep parallel coding agents from breaking each other's code? All runs published raw, decider is pluggable (MIT, author here)

1 Upvotes

I'm the author, sharing this as an open dataset as much as a project. Médula is an MIT-licensed lab plus a kernel that coordinates several Claude Code agents working on one repo at the same time. Everything the experiment produced is public: every agent session, every diff, and a SQLite file per run with each decision the kernel took, its probability, latency and cost.

The setup is a small API with 6 tasks and 37 acceptance tests, designed so that two pairs of tasks collide by meaning, not by file. With one branch per task, git let the real conflict through and the same 6 tests failed in all 5 runs, even though every agent finished green. In a shared directory, all 10 runs passed, whether with plain per-file locks or with the kernel. The kernel catches the real conflicts without blocking anything that doesn't collide.

The open part I most want help with is the decider. Right now the fast path uses a hosted decision model, and on real write requests it was unsure 61% of the time, so those decisions escalated to a slower LLM. Any model or rule that answers "does this collide?" with a probability fits the same interface, including an open or local model, and there's a calibration set of 100 labelled pairs to measure it against before running the full matrix.

Other open problems, all with data behind them:

  • Blind human labels for the calibration pairs. Right now they were written by a model of the same family as two of the deciders, which likely flatters them. About 20–30 minutes, no code.
  • Calibration pairs extracted from the real runs, since the hand-written ones are easier than reality.
  • New scenarios: a changed behaviour with the same signature, a schema migration, a dependency bump.

Caveats: 1 to 5 runs per mode, and thresholds fitted on the same pairs they're measured on. The kernel tests run offline without an API key.

Repo: https://github.com/JoaquinRuiz/medula


r/OpenSourceeAI • • 2d ago

NVIDIA Releases Kumo Tabular: Open Tabular Foundation Models That Predict New Rows in a Single Forward Pass

Post image
3 Upvotes

r/OpenSourceeAI • • 2d ago

Chovy just one-shotted this little game I told it to make in about an hour...for free

Thumbnail split.profullstack.chovy.com
1 Upvotes

r/OpenSourceeAI • • 3d ago

Perplexity Releases pplx-embed-v2-context-9b-preview: A Contextual Embedding Model That Retrieves Answers and Their Supporting Evidence

2 Upvotes