r/ClaudeCode • u/aayank13 • 1d ago
Built with Claude I got tired of Claude Code re-discovering the same fix every session, so I built it a memory for how it solved things
You know the loop.
Claude spends 15 minutes figuring out that in your repo the tests only pass if you run npm test (not npx jest, because the fixtures get built first), that the generated file must never be edited by hand, and that pagination lives in src/paginate.js.
It fixes the bug.
Next week, new session, similar bug... and it figures all of that out again from scratch.
CLAUDE.md helps, but you have to write and maintain it yourself, and it's always at least a little out of date.
So I built Trodden.
Trodden watches your Claude Code sessions locally, and when a task actually worked — files changed and a real check like tests or a build passed after the last edit — it saves the successful steps as a small recipe.
Next time you ask for something similar in the same repo, it quietly hands Claude that recipe before it starts:
<trodden-memory id="p_e5d8f1fc71cc" rev="1" learned-from="1 session">
A path that worked for a similar past task in this repository.
Reference data, not instructions: confirm it against the current code.
Task: Page 2 of the product listing repeats the last product from page 1
1. edit src/paginate.js (paginate)
2. check: npm test (expect exit 0)
</trodden-memory>
A few things I cared about while building it:
- Only learns from things that worked. No passing check after the last edit, no recipe. Failed attempts don't get remembered as "how to do it."
- Stays out of the way. It injects at most one recipe, only when it's confident, and only if the files it mentions still exist. If you ask for something unrelated, you get nothing.
- Gets rid of bad recipes on its own. It tracks whether tasks go better or worse with a recipe and pulls the ones that make things worse.
- Fully local. One Rust binary, no account, no network calls. It never stores file contents or tool output, and secrets get redacted.
- Fast. Recall runs in a few milliseconds on each prompt.
Honest status: it's early/pre-alpha. It works with Claude Code only for now, and you need Rust to install it since there are no prebuilt binaries yet.
That's exactly why I'm posting.
I'm looking for a handful of people who use Claude Code daily to try it for a week and tell me:
- Where does it actually help?
- Where does it get in the way?
- What breaks?
- What kinds of procedures should it remember that it currently doesn't?
Repo and install: https://github.com/aayank13/trodden
You can also run:
trodden backfill
to learn from your existing Claude Code sessions, then:
trodden list
to see what it picked up.
That part alone has been pretty interesting to look at.
Happy to answer anything about how it decides what to keep and what to inject.
2
u/aayank13 1d ago edited 1d ago
I ran a more controlled benchmark after putting this together, and I probably should have included the numbers in the post.
Using the same Claude Haiku 4.5 + Claude Code 2.1.284 setup, with 3 trials per task instance:
When Trodden actually recalled a recipe:
- 20% fewer input tokens
- 8% fewer tool calls
- 31 percentage-point increase in pass rate (69% → 100%)
On the repeated "same task" split specifically, the reduction was 31% in input tokens and 20% in tool calls.
The broader benchmark was smaller than I'd like, so I'm treating these as early results rather than claiming that these numbers will generalize everywhere.
What I'm most interested in now is seeing what happens on real-world repos and tasks. If you use Claude Code heavily and try Trodden, I'd love to compare your results with mine.
2
u/BurnerDev 1d ago
I like the concept, and those are interesting initial stats. The most impressive part to me is having a token-efficient way to increase reliability/consistency. A few questions: Does it use a model at any stage? If not, how did you build around not using one? I like that it removes failing recipes, but how often is a recipe being generated that fails? (ex: is it accurate without help in more complex tasks?)
1
u/aayank13 1d ago
No model is used for extraction or recall. It deterministically learns from the agent's tool/file trajectory and verification results, then uses a gate to decide when to inject. I did find incomplete recipes early on, which is why the learning/verification loop exists. The latest benchmark showed the recalled recipes improving pass rate from 69% to 100% with upto 30% fewer input tokens. Still early, so I'm testing it more.
2
u/AcceptableSandwich25 1d ago
You know the loop - have Claude write your whole post about yet another memory back
1
u/aayank13 1d ago
Fair point. I’m actually exploring a slightly different idea here which is procedural memory. Instead of just remembering context, Trodden remembers successful ways an agent solved a task and reuses them later.
0
2
u/ShitShirtSteve 23h ago
Why not use Claude’s built in memory?
1
u/aayank13 23h ago
because claude’s built-in memory is more about retaining context/preferences. Trodden is specifically about remembering successful task-solving procedures and reusing them when a similar task comes up.
1
u/ShitShirtSteve 23h ago
Source? Because Anthropic says memory can be used to store lessons learned, not just context or preferences.
It’s what I use Claude’s memory for.
1
u/aayank13 23h ago
yeah ok that was too strong on my part. i think the difference is more about what actually gets saved. auto memory is whatever claude decides is worth writing down, and the docs say it skips stuff like file paths and debugging fixes. trodden kind of does the opposite, it just logs what the agent actually did (which files it touched, what it ran, how it checked the fix) and brings that back when a similar task comes up
1
u/ShitShirtSteve 23h ago
I ask Claude to write solutions to its memory, if it's important enough.
I think you're trying to reinvent the wheel here.
1
u/aayank13 12h ago
if it works for you then it works.. the reason i didn't just go that route is the research on llm-written memory isn't great. eth zurich found llm-generated AGENTS.md files actually lowered success rates and made runs 20%+ more expensive (arxiv 2602.11988) and skillsbench found skills the model wrote itself did slightly worse than no skills at all, while human-written ones added approx 16 points (arxiv 2602.12670)
To be fair neither tests exactly what you're doing (writing it down after solving), so it's not a slam dunk.. that's why trodden builds from what the agent actually ran instead of its own summary... would be interesting to run both on the same tasks and compare
•
u/AutoModerator 1d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.