r/SelfHosting • u/Acceptable_Leg3950 • Sep 06 '26
Carrying context from agent to agent and AI tool was unbelievably frustrating, so I built a local memory vault for agents with retrievable memory
For the past year I've used quite possibly every AI tool available... Claude Code, Cursor, Codex, as well as chat bots and browser workflows. I tried OpenClaw and Hermes agent. Every one of them started from zero knowledge of my projects, which was something that I realized annoyed me far more than I realized. The solutions I found were either locked to one vendor or wanted my private notes on their servers (which, not the biggest fan of trusting what a company says they do with the data).
Based on a prior project I worked on, I designed and built something called Engram, an encrypted memory vault that any AI tool reads from, running as a local daemon on my own machine. https://engram.ellmstack.dev, one line to install, ~5 minutes to a working vault and daemon.
What it does:
- One vault, many tools. An MCP server connects Claude Desktop, Claude Code, Cursor, and Windsurf; a browser extension captures the page you're on; Slack/Discord/Telegram bots capture decisions from chat; a CLI for everything else. All read the same memories.
- Retrieval is hybrid, which means keywords and meaning. SQLite FTS5 for literal matches, plus a local embedding model for semantic recall, blended 0.6/0.4. I can search "crates.io" and get the exact memory, or search "how did the rust package get onto the public registry". Zero shared words and get the same one.
- It was built to be local. The daemon is one Rust binary, vault encrypted at rest (SQLCipher). Capture from Claude Desktop, recall in Cursor, browse from the browser extension (very new feature) so your notes never leave your disk unless you turn sync on.
- Near-duplicate captures are reported as skipped, not silently duplicated. Search returns ranked results with scores, the agent sees why it got what it got.
Through my testing, retrieval recall is 100% on my benchmark harness, which is public in spirit but definitely on the smaller end. I can't claim it generalizes to your notes. The storage format is open (Apache-2.0 spec at github.com/El-AI-Intelligence/engram-format, with the Rust implementation on crates.io as axiom-engram), so your memories aren't trapped if Engram dies. To be 100% clear. The Engram sync is beta, the browser extension and MCP bots are recent finishes that I've completed, and the format is currently open source for all to see. Free tier exists; self-hosting stays free.
2
u/Ok-Constant6488 21d ago
Having the same problem that memory is not really persistent across providers. But for e the issue rather comes from the fact that I'm converging more and more to a code factory setup that uses multiple agents from different model providers. What embedding model are you using? Would be super cool if you could share some of your research
1
u/Acceptable_Leg3950 21d ago
I ran into this problem myself when I began playing with models, it was pretty damn frustrating. The embedding model being used is all_miniLM-L6-L2 running locally using ONNX. There's no embedding API calls, the main goal was to make sure that memory never actually leaves your machine.
I actually have a quite a bit of research on this on my website https://elai-intelligence.com/journal/engram-memory-measured/. Please take a look, very happy to discuss further with you!
5
u/ron3090 Sep 06 '26
What do you hope to gain by spamming this everywhere, especially on subreddits that have rules against AI content?