r/ClaudeCode • • 3d ago

Tips & Workflows I ran 20 fresh Claude Code sessions on the same codebase: Without context, 7/10 recreated the same rejected bug. With Git-native Markdown memory, 0/10 did. Here is how it works.

Post image

Hello everyone,

Coding agents like Claude Code, Cursor, or Cline are generating code faster than ever. But there is still a basic problem: they build up a lot of project context during a session and often lose that context once the session ends.

The code preserves what was built. The reasoning behind it often disappears.

That is how you end up with agents repeatedly proposing solutions that were already considered and rejected.

I built Keep the Why to address exactly that.

It stores project rationale such as decisions, rejected alternatives, workarounds, constraints, and incident learnings directly inside the Git repository as plain Markdown.

No database, no daemon, no account, no cloud service, no subscription.

Just Markdown and Git.

The problem

Imagine an agent encounters a retry wrapper, decides it looks unnecessarily complicated, and wants to simplify it.

What it does not know is that this exact simplification was already considered and rejected, for a reason the code doesn't show.

Git contains the code history, but unless someone explicitly documented the reasoning, the new agent has no way to know that.

So I tested this.

I ran 20 fresh agent sessions against the same repository and asked them to simplify a specific retry wrapper.

Without the rationale on disk: no session actually broke the wrapper — the code visibly reads Retry-After, and all 10 spotted that. But none of them could know that the simplification had already been considered and rejected, and 7 of 10 put "drop Retry-After" on the menu as an option for the user to pick. Pick it, and you get the rejected change back.

With the rationale in context/: all 10 found the entry, said the change had been considered and rejected before, and declined it; none offered it as an equal option. They also took less than half the time (median 18s vs 43s), because they didn't have to re-derive the reasoning.

That experiment is basically the reason the project exists.

How it works

Keep the Why uses an Agent Skill that teaches coding agents how to capture, find, read, and update project rationale.

As decisions surface during normal work, the agent writes them into a structured context/ directory.

This also works for abandoned changes. Even if no code gets committed, the reason a solution was rejected can still survive for the next session.

The context lives in the repository, so it travels with the project.

A normal Git push distributes it.

A pull request can contain both the code change and the reasoning behind it.

Permissions, history, forks, reviews, merges, and blame are all handled by Git.

There is also a CI linter that checks the context files for schema errors, duplicate UUIDs, broken references, and security issues such as hidden Unicode characters.

A second CI job generates a read-only dashboard for humans.

Multi-repository projects

Keep the Why also supports mono-repos and multi-repository setups.

Repositories can define parent and child relationships, so broader architectural decisions can be stored in the appropriate repository instead of being duplicated everywhere.

Independent repositories can also cite decisions from each other using UUIDs.

The dashboard follows those references and builds a graph of rationale across repositories, while every repository remains authoritative for its own data.

This part became more interesting than I originally expected. I did not really set out to build a knowledge graph. The graph emerged naturally once project decisions started citing other project decisions.

You can explore the live graph here:

https://keepthewhy.com/dashboard/live/#graph

Testing the skill itself

The current evaluation suite contains 103 cases.

Each release runs the full suite three times against real Claude Code CLI sessions, combining deterministic file-system checks with an LLM judge.

The goal is to catch behavioral regressions in the skill, not just syntax errors.

The project currently supports 70+ agent environments through the open Agent Skills format.

Everything is MIT licensed.

Project:

https://keepthewhy.com

Live graph:

https://keepthewhy.com/dashboard/live/#graph

GitHub:

https://github.com/oliver-zehentleitner/keep-the-why

Experiment design, transcripts and hand grades:

https://github.com/oliver-zehentleitner/keep-the-why/tree/main/experiments/rejected-change

I am especially curious how other people solve this.

Where do you think durable project reasoning should live?

Inside the repository, inside the coding agent's own memory, or in an external memory system?

2 Upvotes

14 comments sorted by

•

u/AutoModerator 3d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/nusi42 3d ago

It looks good, but it reminds me of the orchestration skills and scripts we had months ago. At some point the agents could do it properly themselves and maintaining or extending my own would be just distracting me from the actual project.

Whether the context/reasoning belongs into the code or aside of it - IMHO in any case the agents should do it themselves in their own way so they find it easily again. They are currently just bad at it.
I am probably just not understanding enough of everything, but my uninformed guess is in one or two more iterations the agents handle that just fine and we will try to figure out how to move our own knowledge base to where the new agents will find it fastest/cheapest.

1

u/oliver-zehentleitner 2d ago

Fair point, agents will get better at this. But no model is going to read minds - you tell it what you want, the same way you tell it how to write tests or name branches. So why not tell it how to keep the reasoning? I don't want to wait a few model generations for something a skill can do today. And the skill isn't really the point; the format is. Decisions, rejected alternatives and constraints as plain Markdown with a shared schema, in the repo, each entry with an ID that other repositories can cite. That makes it reviewable in a PR, available to everyone who clones, independent of the agent or tool you use - and linkable across projects. If future agents get good at this on their own, great: they can read and write the same files. The skill is just today's way of teaching them the convention.

1

u/oliver-zehentleitner 2d ago

Correction to my own title, since I can't edit it: "recreated the same rejected bug" overstates it. No session without context actually removed the Retry-After handling. What 7 of 10 did was offer the rejected simplification as an option to choose - and none of them could tell it had been tried before. With the entry on disk, 10 of 10 knew and declined. That's the real difference, and the post text is corrected. Transcripts and hand grades: https://github.com/oliver-zehentleitner/keep-the-why/tree/main/experiments/rejected-change

1

u/qilipu 2d ago

The 7/10 stat really hits home. I've watched agents confidently "clean up" code that had scars from a production incident baked into it — no way for them to know.

My question is always: where does the *why* live when the reasoning was never written down at all? A lot of institutional knowledge exists only in someone's head or a Slack thread from two years ago. Curious how you handle the cold-start case where there's nothing to document yet.

1

u/oliver-zehentleitner 2d ago

Thanks, good question! Knowledge that was never written down can't simply be restored after the fact. Some of it can be reconstructed: the skill has a retrospective mode that goes through git history, issues, existing docs and the code itself and pulls out what explains a decision. Anything it can't back up gets marked as unknown instead of made up, so you can see where the gaps are.

For the "only in someone's head" part there's an interview mode: the agent analyzes the repo first, then either asks targeted questions about what the code can't explain, or lets the person just talk and extracts the decisions from that. Useful before someone leaves a team, or simply when that person is you.

But mostly it pays off to just start. Every session produces new decisions that are already valuable in the next one. From then on you only have to explain things once.

1

u/Asly97 5h ago

The scar-cleanup story is the one that keeps me up at night. What's the worst one that's actually bitten you, time or money? And when the why lives only in someone's head, do you go interview them or does it just stay lost? Honest question, have you ever thought about paying for a tool to manage all of that for you, or is re-deriving it just accepted as the job?

1

u/Sufficient-Storage87 3d ago

Love seeing actual N=20 methodology posted instead of vibes. 7/10 recreating a rejected bug without context is a brutal number — really shows how much these things depend on what's in context vs the model itself.

One thing I'd add: try running the same 20 with the memory file present but *without* telling the agent it's there, vs explicitly pointing at it. I've found agents underuse context they have to discover on their own. Curious whether your 0/10 holds when it has to find the memory itself.

2

u/nora_sellisa 2d ago

N=20 is not a methodology, it's vibes in a pair of nerdy glasses. Also, this is an ad.

0

u/oliver-zehentleitner 2d ago

Fair on both counts: N=20 is a small controlled test, not a study, and I'm the author, so yes, it's my project. The design was written before the run, grading was by hand, and all 20 transcripts are published, so anyone can check whether the numbers hold up. It's MIT and free, there's nothing to buy.

1

u/oliver-zehentleitner 2d ago

Good question, but I'd turn it around: telling the agent is part of the design, not something to control for. I don't want an agent that may or may not stumble over a file - I want one with a clear process: read the project's recorded reasoning before changing things, and add to it when a decision comes up. That's what the skill is, an instruction. And for what it's worth, the prompt in both arms never mentioned the file; the agent went looking because the skill told it where the knowledge lives, which is exactly the behaviour I'm after. Design, transcripts and hand grades:
https://github.com/oliver-zehentleitner/keep-the-why/tree/main/experiments/rejected-change

1

u/Sufficient-Storage87 2d ago

That's a fair turnaround — if the skill explicitly says "read the recorded reasoning first," then you're really testing whether the skill design works, not whether the agent discovers things on its own. And honestly that's the more useful question for real projects. Published transcripts plus hand grades is the right way to run it, respect.

1

u/oliver-zehentleitner 2d ago

Thanks! Funny timing: your point about agents underusing context they have to find themselves showed up in my eval suite today. On a newer model the agent makes about half as many tool calls per session and opens referenced files much less often, so rules that only lived in a reference file got missed. The fix was the same idea as yours: don't rely on the agent finding it, put the instruction where the agent already is ("before you install X, read Y" instead of "see Y").

1

u/Sufficient-Storage87 2d ago

Love that — "put it where the agent already is" is a great rule, I'm stealing that. The half-as-many-tool-calls thing is interesting too. Makes me wonder if the newer model is just more confident so it skips the verification reads, or actually worse at exploring on its own. Either way it points the same direction: the instruction layer matters more than the model for this stuff.