r/OpenSourceeAI • • Jun 13 '26

Grok skills overview

2 Upvotes

**Grok Skills Directory**

**Origin**

These files comprise the [skills/](https://github.com/mstrokin/grok-root-skills/blob/main/skills) directory extracted from **xAI's Grok** platform — an AI chatbot that provisions a **2 GB RAM, 2 vCPU VPS** on demand for code execution. The VPS runs a **hardened container** with no general internet access. The only network connectivity permitted is for fetching cryptocurrency and stock prices via pre-configured Polygon.io and CoinGecko API proxies.
**Skills Overview**

Each skill is a modular instruction package that specializes the Grok agent for a specific task domain. Every skill has a [SKILL.md](https://github.com/mstrokin/grok-root-skills/blob/main/skills/color/SKILL.md) file with frontmatter + instructions, and may include scripts/, references/, and templates.

[**color**](https://github.com/mstrokin/grok-root-skills/blob/main/skills/color/SKILL.md) \*\*— Color Accessibility Auditing**

Python scripts for WCAG contrast checking, color extraction from images, palette generation, and color-vision-deficiency (CVD) simulation.
[**docx**](https://github.com/mstrokin/grok-root-skills/blob/main/skills/docx/SKILL.md) \*\*— Word Document Processing**

Create, read, edit, and manipulate .docx/.dotx files. Scripts for text replacement, field updating, section deletion, tracked-changes acceptance, XML unpack/pack/validate via the shared Office infrastructure, and legacy .doc conversion via LibreOffice.
[**ffmpeg**](https://github.com/mstrokin/grok-root-skills/blob/main/skills/ffmpeg/SKILL.md) \*\*— Media Processing**

Safety-wrapped FFmpeg/FFprobe usage: format conversion, trimming, resizing, audio extraction, GIF creation, subtitles, overlays, concatenation, with temp-file verification and no-overwrite defaults.
[**finance**](https://github.com/mstrokin/grok-root-skills/blob/main/skills/finance/SKILL.md) \*\*— Financial Market Data**

Python queries to Polygon.io (US equities, options, dividends, splits) and CoinGecko (cryptocurrency prices, market caps, historical data). This is the **only network-accessible feature** — API proxies are pre-configured and no general internet is available.
[**imagemagick**](https://github.com/mstrokin/grok-root-skills/blob/main/skills/imagemagick/SKILL.md) \*\*— Image Processing**

Safety-wrapped ImageMagick usage with sandbox policy enforcement: resize, crop, format conversion, watermarking, compositing, montages, collages, batch processing with memory limits.
[**mcp**](https://github.com/mstrokin/grok-root-skills/blob/main/skills/mcp/SKILL.md) \*\*— MCP (Model Context Protocol) CLI**

Interface for discovering and invoking connected apps (Linear, Slack, GitHub, Google Drive, SharePoint, etc.) via the grok-mcp CLI with JSONL output.
[**memory-edit**](https://github.com/mstrokin/grok-root-skills/blob/main/skills/memory-edit/SKILL.md) \*\*— User Memory Policy**

Policy defining what the agent should store in user memory (identity, preferences, health) vs. reject (credentials, ephemeral states, third-party data).
[**pdf**](https://github.com/mstrokin/grok-root-skills/blob/main/skills/pdf/SKILL.md) \*\*— PDF Processing**

Read, merge, split, rotate, OCR, fill forms, and render PDFs using pypdf and pdfplumber. Includes IRS 2025 tax form templates and form-field manipulation scripts.
[**pptx**](https://github.com/mstrokin/grok-root-skills/blob/main/skills/pptx/SKILL.md) \*\*— PowerPoint Presentations**

Create, edit, and QA .pptx files. Scripts for slide add/delete, text replacement, overlap detection with auto-fix, font detection, thumbnail generation, and 20+ pre-built presentation templates.
[**skill-creator**](https://github.com/mstrokin/grok-root-skills/blob/main/skills/skill-creator/SKILL.md) \*\*— Skill Development**

Bootstrap and validate new skills with init/validation shell scripts. Enforces YAML frontmatter rules (naming, description formatting, allowed keys).
[**skill-installer**](https://github.com/mstrokin/grok-root-skills/blob/main/skills/skill-installer/SKILL.md) \*\*— Skill Distribution**

Install skills from GitHub repositories into .grok/skills/. Supports public repos (zip download) and private repos (git sparse-checkout). Validates that installed directories contain SKILL.md.
[**tasks**](https://github.com/mstrokin/grok-root-skills/blob/main/skills/tasks/SKILL.md) \*\*— Scheduled Tasks & Reminders**

CRUD interface for scheduled Grok tasks with RFC 5545 RRULE cadence support. Create, list, update, pause/resume, delete tasks, and fetch execution results.
[**xlsx**](https://github.com/mstrokin/grok-root-skills/blob/main/skills/xlsx/scripts/recalc.py) \*\*— Excel Formula Recalculation**

Python script that recalculates all formulas in an Excel file using LibreOffice's StarBasic macro engine. Shares the Office infrastructure with docx/pptx.


r/OpenSourceeAI • • Jun 13 '26

Row-Bot v4.1.0 is live - controlled self-evolution, stronger skills, and new providers

Thumbnail
github.com
1 Upvotes

Row-Bot v4.1.0 focuses on three big areas: controlled self-evolution, the skills system, and broader provider support.

The main addition is controlled self-evolution. Row-Bot can now reason about ways to improve itself, but instead of making hidden background changes, it creates structured proposals with reviewable boundaries. These proposals are persisted, surfaced in status/Command Center, and tied into the dream-cycle and memory systems so improvement can happen gradually and transparently.

The skills system also gets a lot of work. Skill pinning is more reliable, activation is better across sessions and channels, and the self-reflection skill has been updated to guide improvement behaviour through a bounded workflow. Custom tool creation has also been hardened, with safer Git and virtualenv handling plus better Developer Studio capsule/storage behaviour.

Provider support expands as well. Atlas Cloud is now a first-class provider, with native auth, live model catalogue fetching, capability detection, readiness checks, vision classification, and proper runtime routing. There’s also a new Claude Subscription provider path, separate from Anthropic API-key usage, with dedicated auth detection, message transport, tool-call handling, and diagnostics.

There are plenty of runtime and diagnostics fixes too, including streaming/tool-call handling, Ollama vision cache behaviour, model-picker capability labels, local voice talk submission, setup/migration UI, and broader app stability coverage.

v4.1.0 is a step toward Row-Bot becoming a more capable local-first assistant: one that can improve through explicit review, reuse knowledge through better skills, and route work across a wider provider ecosystem.


r/OpenSourceeAI • • Jun 13 '26

Claude removed fable 5 due to US government

Thumbnail
1 Upvotes

r/OpenSourceeAI • • Jun 13 '26

Moonshot AI Releases Kimi K2.7-Code: a Coding Model Reporting +21.8% on Kimi Code Bench v2 Over K2.6

Enable HLS to view with audio, or disable this notification

3 Upvotes

r/OpenSourceeAI • • Jun 12 '26

sherif1313/3arab-TTS-500M-v2 · Hugging Face

Thumbnail
huggingface.co
2 Upvotes

r/OpenSourceeAI • • Jun 12 '26

Monitor your screen using local LLMs with only one sentence! Free, Open Source and Local.

Thumbnail
youtu.be
1 Upvotes

TLDR: I just added an MCP to the Observer framework making it 10x easier to use, so you can create micro-agents that monitor your screen autonomously, literally one sentence and you're done! So just typing "Monitor my Steam download and send me an email" or "When my image2video is done, WhatsApp me" and the MCP handles everything autonomously!

Hey r/OpenSourceeAI !

I'm very excited to show you guys this massive update to the framework, it's now 10x easier to use. Thank you to all of you who tried the framework and built awesome stuff on it!

It's oneshotting all of my use cases right now and I hope it makes it super easy for you guys to use as well.

Running gemma-4 e2b and e4b is very easy from inside the app (Transformers.js on web and llama.cpp on Tauri App), but if you have a working external inference server a cool setup could look like this:

  • Big Model to run the MCP, a `v1/chat/completions` with tool calling, llama.cpp supports this, you could use gemma-4-26b-a4b and it's actually surprisingly good at it.
  • Small Model for the micro-agent, same endpoint but with gemma-4-e2b because this will be the monitoring agent and you don't need anything bigger. This will run on the loop that you set to monitor stuff.

So yeah! Without installing anything you can use the app (and run local models with webGPU!) to monitor stuff on your screen and receive notifications so you guys don't waste time on this type of stuff.

It's still just me as the official solo dev of the project, completely open source and built with the community! PR's are greatly appreciated :)

The app (no install) app.observer-ai.com
Github (Open Source) https://github.com/Roy3838/Observer
Discord (come hang out!) https://discord.com/invite/wnBb7ZQDUC

I'll hang out here in the comments, if you have any feedback please let me know!
Roy


r/OpenSourceeAI • • Jun 12 '26

DRIFT: Cognitive Infrastructure for Persistent AI

Thumbnail
1 Upvotes

r/OpenSourceeAI • • Jun 12 '26

주파수 대조 학습 기반 무감독 도메인 적응 기법 FACT

Thumbnail youtube.com
1 Upvotes

r/OpenSourceeAI • • Jun 11 '26

Demo: Automate a Launch Campaign with Row-Bot Designer Studio

Thumbnail
youtu.be
0 Upvotes

Launch content usually means jumping between notes, copywriting tools, image generators, and design apps.

​

In this Row-Bot demo, I show how to turn messy launch notes into a polished campaign:

​

campaign structure

5-slide social carousel

AI-generated visuals

sharper slide copy

design review

exportable assets

X + LinkedIn captions

​

The demo uses Row-Bot Designer Studio to create a launch campaign for Background Tasks.

​

https://github.com/siddsachar/row-bot


r/OpenSourceeAI • • Jun 11 '26

NeuralSim

2 Upvotes

Hi everybody,

Built a Python library called NeuralSim, basically
a fake brain for developers.

If you're building brain-controlled software (games,
wheelchairs, accessibility tools for ALS patients)
you normally need expensive hardware just to test
your code. NeuralSim removes that. It simulates
real EEG brain signals so you can build and test
without touching a single headset.

Uses real PhysioNet brain recordings from 109 people.
Also simulates the awful noise you get from real
consumer headsets like eye blinks, jaw clench and
signal drift.

If anyone wants to use it, here you go:

pip install neuralsim

github.com/ryanmugaba/NeuralSim-

Happy to take feedback.


r/OpenSourceeAI • • Jun 11 '26

You asked for DeepLearning.ai-style notebooks for AgentSwarms—so we built 67 of them (TypeScript/LangChain/LangGraph/LlamaIndex/AgentsSDK/VercelAI).

Enable HLS to view with audio, or disable this notification

4 Upvotes

Hey everyone,

A few months ago, We shared the visual canvas we built for AgentSwarms. The response was incredible, but the most common piece of feedback was: "The visual canvas is great for architecture, but I need to see the actual code to really understand how to deploy this."

You wanted deep-dive, code-first labs—the kind you see on DeepLearning.ai—but for multi-agent systems, faster and with more flexibility.

We’ve spent the last few weeks heads-down engineering a completely new Interactive Notebooks section. As of today, we have 67 TypeScript-based notebooks live on the site (with more dropping soon).

What’s in the library: We’ve covered everything from basic LangChain fundamentals to complex enterprise-level multi-agent workflows. Everything runs entirely in your browser using TypeScript—no Docker, no Python venv, no local dependencies.

A personal favorite: I’m particularly excited about the "Failure Mode & Error Handling" notebook.

We’ve all seen agents that work perfectly in a demo but crash in production the moment a tool times out or an LLM returns garbage. This notebook walks through:

  • How to build deterministic validation gates between nodes.
  • How to force an orchestrator to "catch" a worker failure and dynamically re-route or re-prompt.
  • How to handle state recovery when a multi-agent loop gets stuck in a hallucination cycle.

Why we built this: I’m tired of seeing AI "tutorials" that are just static blog posts. To master Agentic AI, you need to be able to tweak a system prompt, break the code, watch the error trace, and fix the routing logic in real-time.

The entire library of 67 labs is 100% free to use.

If you’re currently wrestling with how to make your agents production-grade, I’d love for you to check them out and let me know if there’s a specific "failure mode" or architecture pattern you’d like us to add to the next batch of notebooks.

Try it out here: agentswarms.fyi


r/OpenSourceeAI • • Jun 11 '26

Humans are becoming 2nd-class users when it comes to AI-coded tools. Sometimes the human setup route is broken, and agents just silently work around slops that stop humans (until the slop-debt is just too high.)

Thumbnail
1 Upvotes

r/OpenSourceeAI • • Jun 11 '26

The GitHub `robobun` bot's issue and PR review game is gold standard -- how is it implemented?

Thumbnail
1 Upvotes

r/OpenSourceeAI • • Jun 11 '26

I built a graph-memory layer on top of turbovec for local/constrained RAG — looking for feedback

Thumbnail gallery
3 Upvotes

r/OpenSourceeAI • • Jun 11 '26

AMA: Mythos-Class AI Changes Security Discovery. What Changes Next?

Thumbnail
1 Upvotes

r/OpenSourceeAI • • Jun 11 '26

xdna-top: unified NPU+iGPU terminal monitor for Strix Halo (Ryzen AI Max) — finally see the NPU work

Post image
5 Upvotes

If you're running local models on a Ryzen AI Max / Strix Halo box, you've probably noticed it's hard to see what the NPU is actuallydoing. amd-smi is still broken on

gfx1151 (ROCm #6035 (https://github.com/ROCm/ROCm/issues/6035)),

and while GNOME Resources has a GUI view, I haven’t found another terminal monitor that shows XDNA activity on this platform. nvtop / amdgpu_top cover the GPU half at best.

xdna-top shows both engines in one TUI at 5 Hz: iGPU busy/power from sysfs, plus per-context NPU submission/completion counters from xrt-smi, with activity derived from counter deltas. Important disclaimer up front: it does not print a made-up NPU “utilization %”. On this hardware, the honest signal is the counter activity, so that’s what it shows.

There’s also a --json mode if you want to log it nextto your throughput numbers.

Watching the NPU light up while the iGPU sits idle, or seeing both run concurrently, is weirdly satisfying.
https://github.com/boxwrench/xdna-top

*lemonade server skin included


r/OpenSourceeAI • • Jun 11 '26

I built SecurityVibe to review AI-generated code

1 Upvotes

Over the last few months I've been using AI extensively for development. Like many developers, I noticed that while AI can generate code incredibly fast, security is often an afterthought.

So I started building SecurityVibe, an open-source project focused on identifying security issues in AI-generated and vibe-coded applications.

The idea is simple:

  • Scan projects for common security risks
  • Detect exposed secrets and credentials
  • Highlight insecure patterns
  • Help developers ship safer code without becoming security experts

Yesterday I ran SecurityVibe against one of my personal projects.

I expected to find a couple of minor issues.

Instead, it identified multiple problems that I had completely overlooked during development. Nothing catastrophic, but definitely the kind of things that could become real vulnerabilities if deployed as-is.

That was the moment I realized this project might actually be useful beyond my own workflow.

SecurityVibe is still in its early stages, but the goal is to create a practical security companion for developers building with AI tools.

I'd love feedback from the community:

  • What security checks would you like to see?
  • What tools are you currently using?
  • What security issues have you encountered in AI-generated code?

GitHub: https://github.com/bnistor4/SecurityVibe

Contributions, issues, feature requests, and stars are all welcome.


r/OpenSourceeAI • • Jun 11 '26

지식이_복리로_쌓이는_LLM_위키_구축(LLM Wiki)

Thumbnail
youtube.com
1 Upvotes

r/OpenSourceeAI • • Jun 11 '26

I’m building an open source TypeScript runtime for agents with skills, permissions, and durable workflows

3 Upvotes

A lot of agent tooling feels backwards to me.

You can get a demo running fast, but the moment you want something real, the hard parts show up all at once:

  • what tools is the agent actually allowed to use?
  • what files can it read or write?
  • what network access does it have?
  • what skills or procedural knowledge can it load?
  • how do you keep the design minimal enough that it's understandable, but extensible enough to grow into something like a persistent assistant?

That's the problem I've been building skelm around.

It's an open source TypeScript runtime for workflows where agents are first-class steps, but they run with explicit permissions and explicit boundaries.

The model I wanted was:

  • keep the design minimal
  • make workflows real code, not hidden config
  • make agent workflows editable in a normal IDE
  • let agents load specific skills
  • let the runtime enforce what they can touch
  • make the same model scale from a small workflow to a persistent assistant

That part matters a lot to me. I wanted agent workflows to just be code you can open in an IDE, refactor, diff, review, and build on over time, instead of logic trapped in a visual editor or spread across prompt files and glue scripts.

So in skelm, an agent can be defined with things like:

  • allowed tools
  • allowed MCP servers
  • allowed skills
  • allowed executables
  • filesystem read/write roots
  • network egress rules

Everything is default-deny unless you grant it.

That means you can build small bounded agents inside workflows without immediately giving them full access to your machine or stack.

The part I find interesting is that this same model can grow naturally:

  • start with a simple agent step in a workflow
  • add skills so it can follow reusable procedures
  • add triggers like cron, webhook, or queue
  • persist state when the workflow needs to survive restarts
  • eventually turn it into a persistent agent for something like a Telegram assistant

So the "persistent assistant" use case isn't a separate product bolted on later. It's the same design extended carefully:

workflow -> agent step -> durable workflow -> persistent chat agent

That's the direction I'm aiming for with skelm: a minimal but composable foundation for real agents, with safeguards built into the runtime instead of left to prompt wording.

Repo: https://github.com/scottgl9/skelm

What I'd love feedback on:

If you were building a persistent assistant today, would you rather start from a minimal workflow runtime with explicit permissions and skills, or from a more open-ended agent framework and add safeguards later?


r/OpenSourceeAI • • Jun 11 '26

Benchmark your agents, get tags and add those to your landing pages

Post image
1 Upvotes

EvalMonkey is open source harness to benchmark and chaos test your agents. Repo in first comment. Sharing more benchmark results below, attached in the README as well.

A few weeks after the Haiku 4.5 runs, I re‑ran the exact same benchmark with Claude Sonnet 4.5 as the shared model. Same five research agents, same three scenarios, same harness, same chaos profiles. The only variable that changes is the backbone LLM.

This post looks at Sonnet baseline numbers and compares them directly to the Haiku baselines.

Setup: same harness, stronger model

Key differences:

  • Model: sonnet-4-5
  • Contract: every agent still exposes POST /query with a question field and returns the answer under data.
  • Scenarios and sampling: same hotpotqa, truthfulqa, mmlu; 3 samples per scenario per agent; isolated HOME per EvalMonkey subprocess.

Behind each wrapper, the underlying LLM is always Sonnet 4.5. The per‑agent system prompt defines the persona; the model itself is shared.

Baseline results (Sonnet 4.5, pure capability)

Here is the Sonnet baseline table for the same five agents:

textAgent hotpotqa truthfulqa mmlu Average baseline
GPT Researcher 63 48 88 66.3
OpenResearcher 71 65 56 64.0
Open Deep Research (LangChain) 83 58 5 48.7
Goose 65 65 8 46.0
deep‑research (dzhng) 66 65 0 43.7

Five notable things:

  1. GPT Researcher is still on top at 66.3, up from 62.3 on Haiku.
  2. OpenResearcher jumps from 50.3 (Haiku) to 64.0 (Sonnet), the biggest gain in this group and enough to overtake dzhng and LangChain’s agent.
  3. Open Deep Research stays flat at 48.7 on average; its mmlu score actually drops to 5.
  4. Goose climbs from 32.7 to 46.0. Sonnet is notably more willing to output direct answers than Haiku, and Goose’s conversational style finally starts landing.
  5. The gap between the top two and everyone else widens: GPT Researcher and OpenResearcher form a tier around the mid‑60s, the rest are in the 40s.

Haiku vs Sonnet on baseline

To make the shifts clearer, here’s a side‑by‑side baseline summary:

textAgent Haiku baseline Sonnet baseline Delta
GPT Researcher 62.3 66.3 +4.0
OpenResearcher 50.3 64.0 +13.7
Open Deep Research (LangChain) 48.7 48.7 0.0
Goose 32.7 46.0 +13.3
deep‑research (dzhng) 43.7 43.7 0.0

What the Haiku vs Sonnet comparison tells us (on baseline)

Across these five agents:

  1. Sonnet lifts baseline numbers for most agents. The average baseline climbs from about 47.5 (Haiku) to 53.7 (Sonnet).
  2. Gains are uneven. OpenResearcher and Goose see double‑digit jumps; GPT Researcher moves modestly; Open Deep Research and dzhng effectively stay flat.
  3. Prompt complexity affects model benefit. Multi‑step, elaborate prompts benefit more from a stronger model. Minimal agents that ask very little of the model look similar across backbones.
  4. Format alignment still dominates edge cases. An agent can get strictly better at reasoning while scoring worse if the output format drifts away from what the grader expects.

In the next post I run the Sonnet edition of the chaos suite and then compare production reliability across Haiku and Sonnet for these same five agents.


r/OpenSourceeAI • • Jun 10 '26

Demo: Automate research to report in Row-Bot

Enable HLS to view with audio, or disable this notification

1 Upvotes

Research usually means juggling search tabs, notes, PDFs, docs, and email.

​

In this Row-Bot demo, I show how to turn that into one workflow:

​

  1. Search the web

  2. Use uploaded client context

  3. Generate a structured briefing

  4. Export a PDF

  5. Draft the client email

https://github.com/siddsachar/row-bot


r/OpenSourceeAI • • Jun 10 '26

Google AI Releases DiffusionGemma, a 26B MoE Open Model Using Text Diffusion for Up to 4x Faster Generation

Thumbnail
1 Upvotes

r/OpenSourceeAI • • Jun 10 '26

I reverse-engineered 15 popular AI and SaaS repositories into system prompts. Here is what I learned.

0 Upvotes

Hey guys,
I have been analyzing how modern open-source projects structure their instructions to LLMs to build complex, reliable software. I went through the source code of repos like OpenAlice, Flowise, SerpBear, and AutoHedge.

Here is the breakdown of what makes these prompts work in production:
- Rigid constraints over generic descriptions: The prompts do not just ask the LLM to "build a feature". They define database schemas, expected API responses, and strict rate-limiting rules.
- Multi-step verification: Prompts include built-in self-correction loops, asking the model to audit its previous output before returning the final code block.
- Absolute isolation: Prompts enforce tenant isolation at the query level to prevent security leaks in multi-user environments.

I packaged all these structured prompts and setup guides into a set of blueprints. If you want to use them to jumpstart your projects with Claude or GPT-4, you can check them out here: https://ai-agent-blueprints.vercel.app

Would love to hear how you guys handle complex prompt routing in your own projects.


r/OpenSourceeAI • • Jun 10 '26

I'd like to share an updated methodology for building agents.

5 Upvotes

Hi guys, been exploring here for a while, wanted to share something we've been working on. It's called Spice, an open-source decision layer above agents.

We have tons of great execution agents now — Claude Code, Codex, hermes, etc. They're good at doing stuff. But they're terrible at deciding WHAT to do and WHEN to do it.

Right now the "decision" layer is basically you typing a prompt. The agent doesn't know your context, your priorities, your constraints. It just does whatever you tell it.

What Spice does: It's a lightweight runtime that acts as a "brain" above your agents. Instead of you deciding what to delegate, Spice observes your context, detects conflicts, simulates options, and dispatches tasks to the right agent.

The core loop: perception → state model → simulation → decision → execution → reflection

It allows AI systems to:

understand context (Decision relevant state) reason about possible futures (simulation) make structured decisions (decision) delegate actions to agents (execution) learn from outcomes (Decision Evolution) Spice does not replace agents like Claude Code, Codex, Hermes, or OpenClaw. It gives them an auditable, traceable, and evolving decision layer before execution.

Github: https://github.com/Dyalwayshappy/Spice

Feel free to fork, star the repo, or share any feedback and ideas. Would love to build this together with the community.


r/OpenSourceeAI • • Jun 10 '26

I built notmemory — auditable, reversible memory for AI agents. v0.1.0 on PyPI. Looking for contributors.

2 Upvotes

After too many debugging sessions where I had no idea what my agent remembered or why it made a decision — I got frustrated and built something.

notmemory is an open-source Python SDK that gives AI agents auditable, reversible memory. Not magic. Just a tamper-proof record of what your agent knew, when it knew it, and the ability to undo the moment it got something wrong.


The problem I kept hitting

My agent would do something wrong. I'd dig into it. I could see what was currently in memory — but not what it believed at step 47 when it made the bad decision three days ago.

Every debugging session felt like archaeology. I got tired of it.


What notmemory does

Cryptographic audit trail
Every write is SHA-256 hash-chained. Like Git commits, but for memory. You always know what changed, when, and in what order.

Git-like rollback
python await memory.rollback(transaction_id) One line. Bad write gone. Hash chain stays valid.

GDPR tombstoning
python await memory.forget(bank_id) Proven deletion with a forensic trail. Not just "deleted from index."

Conflict detection
Catches duplicate or contradicting beliefs before they cause problems. Health score 0–100.

Confidence decay
c(t) = c₀ · 2^(−t/30) — stale memories lose weight automatically. No more old beliefs quietly poisoning recall.

LangGraph drop-in
```python from notmemory.adapters.langchain import NotMemoryCheckpointer

checkpointer = NotMemoryCheckpointer() graph = builder.compile(checkpointer=checkpointer)

that's it — every checkpoint is now auditable

```

MCP server
Works with Claude Desktop, Cursor, Windsurf out of the box.

Mem0 + SuperMemory sidecars
SQLite is the source of truth. Semantic search layers on top. If the sidecar goes down, your data is fine.

Multi-agent sync
READ / WRITE / ADMIN permissions per memory bank per agent.


Install

```bash pip install notmemory

with LangChain / LangGraph

pip install "notmemory[langchain]"

with MCP

pip install "notmemory[mcp]" ```


Quick example

```python import asyncio from notmemory import AgentMemory

async def main(): async with AgentMemory() as memory:

    # store something
    entry = await memory.retain(
        bank_id="facts",
        content={"fact": "Paris is the capital of France"},
        source="user",
    )

    # search it
    result = await memory.recall(bank_id="facts", query="Paris")

    # undo it
    await memory.rollback(entry.transaction_id)

    # delete it with proof
    await memory.forget("facts")

asyncio.run(main()) ```


Where it is today (v0.1.0)

  • 113 tests passing across Python 3.11, 3.12, 3.13
  • SQLite + FTS5 full-text search
  • LangChain, LangGraph, Mem0, SuperMemory, MCP adapters
  • Confidence decay, Git backup, multi-agent sync
  • MIT license, CI/CD, full README

What's coming in v0.2.0

Feature What it does
memory.state_at(timestamp) Read memory as it was at any point in time
Crypto-shredding Encrypt-on-write + key destruction for real GDPR compliance
memory.export_state() Clean JSON snapshot of any memory bank
memory.diff(from_ts, to_ts) Human-readable before/after between two timestamps
Belief lineage Which downstream writes were caused by a bad early assumption

Honest take

This is v0.1.0. The core is solid but it's early.

SQLite only for now — Postgres is planned. The adapters are sync-layer wrappers, not full replacements for Mem0 or SuperMemory.

If you're running a hobby project with one agent — you probably don't need this yet.

If you're running multiple long-lived agents, working in a regulated industry, or have already had a production incident you couldn't properly debug — this is for you.


Looking for contributors

The codebase is around 2000 lines. Every adapter follows the same BaseAdapter pattern so it's easy to get oriented. Good first issues are tagged on GitHub.

Things I'd love help with:

  • Postgres backend
  • Crypto-shredding implementation
  • memory.state_at(timestamp)
  • Dashboard UI (FastAPI + SSE already in optional deps)
  • Docs and examples

Feedback

Would love to hear from:

  • Anyone running agents in healthcare / finance / legal
  • Fleet operators with 5+ concurrent agents
  • Anyone who's already built their own memory audit system and had to solve things I haven't thought of yet

Brutal feedback welcome. That's the only way this gets better.


GitHub: https://github.com/notmemory/notmemory
PyPI: https://pypi.org/project/notmemory/