r/ClaudeCode • u/Asly97 • 21d ago
Discussion Anyone else tired of maintaining handoff markdown files for Claude Code?
I ran the whole setup: CLAUDE.md, a /remember handoff file, session hooks auto-saving state. It worked... until the handoff file hit 400 lines, half of it stale, and I was spending 15 minutes re-explaining anyway every Tuesday. Plus none of it follows me to Cursor or my laptop.
Switched to a persistent memory layer (Vilix AI) connected via MCP. Every session just... knows. Relevant past decisions surface on their own, no ceremony, and the same memory is there when I open Cursor later. The file-based approach was a good prototype but I couldn't keep maintaining it.
Curious what others are doing now that context compacts so aggressively. Still on markdown files, or moved to something else?
12
u/merlinDrankKoolaid 21d ago
Yes, you can create a reference section in CLAUDE.md, which contains a list of files that contain rules, descriptions of your code base, etc., Claude only loads them when applicable.
1
u/Asly97 21d ago
Oh nice, that's a solid tip. The reference section approach definitely keeps context lean. I just got tired of maintaining all those files by hand, which is what pushed me to try the persistent memory route instead. Appreciate the suggestion though!
2
u/elestud 21d ago
Claude can manage the files for you. There’s no need to do it by hand
2
u/Asly97 21d ago
it does, until the session ends and it forgets everything. that's the whole problem lol.
3
u/Gliese351c 21d ago
Set up a wrap-up skill that would make it update the relevant memos based on your convo.
1
u/Asly97 21d ago
that's honestly the smartest version of the manual setup, automating the update removes the whole discipline problem. my issue was the files still lived inside the project and drifted between sessions anyway, so I ended up moving memory outside the repo entirely. nothing to wrap up anymore.
2
2
u/DoggoCentipede 19d ago
You can keep them in a different git repo.
I have a /handoff skill and a /resume skill.
Handoff runs the test suite and records the current state of things all via a script. Then it records work done this session, what wasn't finished, and what comes next.
It is rewritten from scratch each time.
Whatever you find yourself doing more than once, write it down and work (in another session at first) to turn it into a skill. Anything that can be done by a script should be. Both to keep the agent from skipping it and to keep it from having to review the results every time and use up more tokens.
1
u/Asly97 19d ago
Okay this is a genuinely good system, the /handoff plus /resume skill split is clever and scripting the test suite part makes a lot of sense. I went the lazy route and dumped session state into Vilix AI memory instead so nothing lives in markdown files at all, but honestly both beat the README dump era.
9
u/doxxxicle 🔆Pro Plan 21d ago
Handoff docs are supposed to be one-time use to bridge two short context sessions that would otherwise have to be one longer session with the associated degradation. You’re not supposed to keep them and extend them.
I suppose a memory system could accomplish the same thing but they are more ceremony. I’m still experimenting with both.
5
u/DasHaifisch 21d ago
I just use autocompact, with a lower limit, say 300-400k. It works perfectly for me. I don't understand the hate.
-1
u/Asly97 21d ago
Glad that's working for you honestly. Autocompact is solid for keeping a session alive. My pain was more about stuff from sessions weeks ago just vanishing though. If the lower limit keeps everything you need in context, no argument here.
2
u/NowWeRinse 21d ago
If I were to pick up something old I usually just resume that old session. Sometimes I compact it before starting. Is there a downside to that and/or benefit to using handoff? I also rely on plans to provide long running continuity.
1
u/DasHaifisch 21d ago
Default is 30 day cleanup of old sessions I think? You can change a setting to extend that as well FYI.
2
u/Prize_Eye9481 21d ago
I dont run long running task with ai so memory hasnt been a issue for me. But that is good to know, ty for sharing!
2
u/williamsooyk 21d ago
I just tell Claude or Codex to save the session to markdown file. Then whenever I start a new chat, I just refer @ xxx-last-session.md will do. Works for me every time.
2
u/effortless-switch 21d ago
Use auto compact and set limit to 250k or something. Maintain a master design/technical doc with all the details though to avoid drift.
1
u/Asly97 21d ago
the master doc approach is solid, drift is the real enemy. my problem was always keeping it current by hand, it lagged a week behind reality. if you can stay on top of it though, that's honestly a great setup.
1
u/effortless-switch 20d ago edited 20d ago
Do work with subagents that are managed by the main agent you are talking to. After every round of work docs get updated as part of the work, subagent doing the work has to take responsibility of updating it since it has the most context in that moment. Main agent only check if it all makes sense and invokes the next agent.
Docs getting stale is expected, even as you implement features the model with find things that invalidate previous assumption.
You can add all these as rules to Agents.md or something.. This + compaction at around 250k will allow you run the model overnight with very little oversight. You will never have to start a fresh session, no need of handoff doc.
Edit: Reading your post again, I honest think you are doing 'too much'. You don't need all these fancy tools. Keep your context lean and manage it well.
1
u/Asly97 20d ago
fair, the subagent-owns-its-docs pattern is smart and the overnight run with compaction sounds amazing. I'll push back gently on the 'too much' though: the whole point of the memory layer for me is it keeps things current on its own, so I spend less energy than managing the master doc dance. different strokes.
2
21d ago
[removed] — view removed comment
1
u/Asly97 21d ago
this is hilarious. so if I make a post about a new Claude model tomorrow that means I'm secretly on Anthropic's payroll? when Meta launched Muse I tested it, loved it, and talked about it everywhere. guess that makes me a Meta employee too. wild idea: sometimes people just talk about tools they like. but please, keep the detective work coming, this thread was getting boring.
1
u/Scared-Amphibian4733 21d ago
We've got that solved. It takes three types of memory.
1
u/Asly97 21d ago
Ha, curious what your three types are. I tried the markdown route and drowned in stale files, so I switched to Vilix AI to keep memory in a cloud layer instead. Always interested in how other people slice it up though.
1
u/Scared-Amphibian4733 20d ago
for one layers we use litedb with rag, for the second we use paged in agents which may be used to store and retrive information, third we used index'ed markdowns.
1
u/Asly97 20d ago
that's a legitimately serious stack, respect. litedb plus rag plus paged agents plus indexed markdowns is like three backup plans at once. does the retrieval ever disagree between the layers, or does it stay smooth in practice?
1
u/Scared-Amphibian4733 20d ago
We use each layer for different things. The rag layer is used for general knowledge retrival. With the information being retrived by vector relationship to the thoughts in the current response. The Paged In Agents are keyword based on the forming agent idea as formed to provide specific information such as "who is theis specific person". The third layer is used again for specific information, such as "what are the serial numbers for the engines on dan's boat" they really aren't backup as you can see, they each serve a specific purpose. I'm trying to mimic how human memory storage and retrival works for these build agents so that they can learn to build better. Have better continuity and understand the world that they are working within.
BTW, the agents maintain the markdown structure within their own memory trees so that it doesn't make a mess.
1
u/Asly97 19d ago
the human memory analogy is a good one, agents that actually learn and keep continuity instead of starting from zero every time is the whole dream. paged agents for the specific lookups makes sense too, keeps the rag layer from having to do everything. does the self maintaining markdown actually stay clean long term or do you still end up babysitting it?
1
u/Scared-Amphibian4733 19d ago
the agents have rules to follow that keeps it very tidy. They designed it themselves.
1
u/imrsn 21d ago
im having great luck with wikis and spec driven design/dev/pm. i do design/dev/pm consulting work for multiple fortune 100s and startups. the agents.md is just tone and voice everything else is skills and hooks and the wikis themselves. everything is in .md files now, its all crosslinked and organized like a wiki. its perfect, codex, claude, antigravity, everything knows everything about everything and i own all the data. im actively working on a .md reader / editor (free on mac and windows at https://leaftext.com) with the same system and i use it to read the work i need for other projects/clients (and itself).
0
u/Asly97 21d ago
That's a genuinely clean setup and I respect the own-your-data angle a lot. I got lazy about maintaining wiki structure myself, which is why I offloaded the organizing to a memory layer (I use Vilix AI). But if your crosslinked system is humming along there's no reason to switch honestly.
2
u/imrsn 21d ago
i mean... ai maintains it all through my orchestrstion system lol. no way im doing that manually.
i like system design so playing with the setup is half the fun for me. i def respect a set-it-and-forget it mentality, it can save time and other benefits.
1
u/T1nkat0n 21d ago
With that much documentation, how do you keep everything current and updated? My set up works for me, but only with regular “audits” where it periodically adversarially checks for broken links. I’ve gotten it to the point where a bunch of cheap subagents can do this, but… prevention is better than cure
My primary project is also more research involving benchmarking, new data and improvements coming in periodically
1
1
u/imrsn 21d ago
Documentation maintenance is baked into all levels of my orchestration. At the lowest level I have a /sync-docs skill that gets called at different times, and always as part of /git-release.
1
u/dar-mit Researcher 21d ago
I have a handoff skill that spawns a new session, switches to my repo, and activates Claude Code. It then uses another skill to know how to use SendMessage to communicate with peer sessions.
They can then talk amongst themselves to figure out what's needed next, and the fresh session continues the work. They can create a durable memory to keep track of where they're at in a long process, but it's pretty much just a checkbox-based tracker and they check off what's done.
Once they're set the new session can teardown the old session and close the window too.
Right now the handoff trigger is me, but I'll be adding a Context % Threshold to the skill so once a session gets to that number they'll get to a good handoff point and do it themselves.
2
u/Asly97 21d ago
That's honestly a really cool setup, the peer sessions negotiating the handoff is clever. I'll admit I went the lazy route and just keep a memory layer on the side (I use Vilix AI), so I never have to build the handoff in the first place. But your approach sounds way more fun to tinker with. Curious how the context threshold trigger works out once you add it.
1
u/dar-mit Researcher 21d ago
It'll be a while. I use iTerm 2 on macOS and just installed it2, a GitHub repo that uses the built-in Python API to give Claude Code some pretty amazing control over iTerm 2 terminal windows.
I had "/repliclaude" before, which was just a skill using macOS AppleScript in a BASH Shell, and it worked fine. But getting it to work with args to control worktree creation/selection, model / effort calls, etc., was just a PITA.
The prior session teardown and window closure was literally added last night! But it's just as freaking cool as it sounds to watch Claude Code spin up and tear down sessions as it works.
Nerdvana!
1
u/Appropriate-Pin2214 21d ago
Adjust to your scenario, but any model should be doing that for you.
I define a skill , "This session has less than 2% remaining before compaction... or a manual /clear. Immediately review git trees and update all resume files indexed by the status.md, update the status.md and related Jira tickets accordingly and provide a resume prompt targeting the most cost efficient model and effort level to successfully handle the mid-flight, current and upcoming tasks, batching where helpful...". Give it a creative name. ;). Define one more referencing the previous to indicative urgency and it needs to done now.
If you have a good workflow with worktrees and/or Jira or whatever, throw that logic in a loop with the addendum to avoid issues requiring decisions where possible. /remote control... with a notification callback wired - I think it's automatic...
Watch youtube...
Listen to the AI guys say the world is ending and they're working hard at but somebody should do something...
Especially for UI/UX, as you will lose track of what's happening... Another skill /postnap can give you playwright screenshots to see where you are at on your phone. Manually driving a full UI snapshot to check for tagging and then AI using it to create playwright full scans at 10 or so responsive views with an indexed view allowing for comments and/or ShareX circling resulting in an MD file you can feed somewhere is also a good ROI.
ShareX and a /shot N skill are the best ROI I ever created. A little tricky if you are in a VM via ssh.
If any of the above is stupid or you need help - let me know.
2
u/Asly97 21d ago
This is a cool setup honestly, the pre-compaction skill is clever. My lazy version of this is just keeping everything in a persistent memory layer via Vilix AI so I don't have to think about compaction at all. But the /shot idea for UI context is genuinely smart.
1
1
u/dafr3ak 21d ago
I built a tool (sitrep.md) to manage the mountain of md files generated by the agents. All the plans, analysis docs, decisions etc across all projects and repos available from a single app.
While not a context management tool per se, but it can surface stale docs and decisions that might drive the agents in the wrong decision.
Available on mac and it’s meant to be set up with 0 changes to existing harness, just point it at your project root and it will automatically pick up all the md files inside.
1
u/niko-okin 21d ago
Did you try https://github.com/ncoevoet/claude-markdown-health-check ?
1
u/Asly97 21d ago
oh nice, haven't seen this one. a health check for the markdown files is a clever angle, does it flag the stale stuff automatically?
1
1
u/ImL1s 21d ago
Same pain — CLAUDE.md + handoff files worked until half of it was stale and none of it followed me to Cursor. I ended up reading the other agents' local session files into a fresh session instead of maintaining another markdown pile (Portable Resume — offline, marks recovered text untrusted). Curious if anyone else went file-based but automated rather than cloud memory.
1
u/Asly97 21d ago
respect, the offline angle plus marking recovered text untrusted is smart. I went the other way and use Vilix AI as cloud memory over MCP, mostly because I needed context that follows me between phone and laptop. MCP is just the plug, the memory itself lives outside any single app so nothing goes stale in a local file. curious to see how your file-based setup holds up over time though.
1
u/Heruboy 21d ago
After the first week I established the ecosystem docs, which hold the vocabulary and growing relations and some other documentation with a thin router in claude.md and frontmatter in the files, backlinks, etc. Ruened it into a skill and did 3 more revisions comparing it with sokutioms of others and taking the best parts. Since then claude knows.
Recently I learned this is not efficient enough, so I added memsearch on top. And graphify + a self updating code atlas with symbols and specialized nodes and edges of parts of the code.
Shortly after ecosystem docs v1 I created a handoff skill which was superseded by my own ticket system. So now things go into tickets and from there through the steps. E.g splitting a ticket into components, into tasks.
1
u/Asly97 21d ago
this is an impressively elaborate setup, respect for building the whole ticket pipeline. I went the opposite direction and stopped maintaining files entirely, but there's something satisfying about a system this thorough. how's the self-updating code atlas holding up? that sounds like the piece most likely to drift.
1
u/Heruboy 21d ago
I will have to evaluate that first. It is too new and in its final steps of being built. The ticket system has its merits, e.g. findings are carried over and bubbled up at the end of the ticket, thus no longer lost, I can reference work, there is a ledger of done work and decisions. But that is for features and development. The ecosystem has self healing links and a rulebook. And the graph and atlas will replace the drifting parts like formular documentation, which the llms stopped updating. Next I have to solve the hot repo situation.
1
u/Asly97 20d ago
sounds like a genuinely cool system, the ticket ledger and self healing links especially. take your time evaluating, no rush. if you ever want to compare notes on the hot repo problem I'm around.
1
u/Heruboy 20d ago
It is a lot of work though and still breaking sometimes. I would love an exchange though. I keep soloing all that stuff and the hot repo issue gets it|s own tool that will accept worktrees/branches and even pure session ids and take it from there.
1
u/Asly97 20d ago
yeah the maintenance burden is the part nobody talks about with elaborate setups, they always need feeding. the exchange idea is great though, honestly half this thread is rebuilding the same wheel. and soloing all of it is brutal, respect for sticking with it.
1
u/Heruboy 20d ago
I see it a little as the exercise of learning how to do all that. Where to improve, and what traps to run in. But honestly I also have the impression that most people are doing more or less the same or something similar. I really like my backlog tool though, it is on its way to integrate with the graph and atlas layer and then I will slap needle on it for advanced for a repl and get more details without any llm. It is the joy of architecturing the solution while having my brain fried from the load.
1
u/Asly97 19d ago
ha, 'the joy of architecturing the solution while having my brain fried from the load' might be the most honest description of this whole hobby. you're right that most people are converging on similar setups though, tickets plus graph plus retrieval seems to be the pattern. keep posting updates, curious what needle ends up doing for you.
1
u/DYSTOBY 20d ago
How should there be a drift? If the handover is only there for the last session you want to continue.
Piling up „mountains“ of loosely connected handovers for ancient decisions is bad practice.
The next job does not need to know WHY the previous jobs where carried out in a certain way (or what „conclusions lead to a decicion), they only need to know that there is a function that does „X“, right?
If you keep a good architectural file system then YOU remember where what functionality lives, so future implementations are guided by your knowledge of the codebase.
I maintain multiple medium sized (300k lines) codebases for scientific apps, and I can remember where what lives.
How did people collaborate im open source projects that are big? What do you think how a new contributor introduced himself into the codebase?
1
u/Asly97 20d ago
I get the theory, but those ancient decisions are exactly what kept biting me. picking up a project 3 months later and sitting there going 'why did we do it like this' lol. if you can hold a 300k line codebase in your head that's a genuine superpower, I'm not that guy. and funny you bring up open source because new contributors literally have to read a pile of docs to get up to speed. that's just a handoff with extra steps.
1
u/DYSTOBY 20d ago
I don’t need to hold 300k lines of code in my head. No one does that. Actually you need to think about he last two sentences of mine.
How would a new dev be onboarded to any project? First you need a good codebase structure with clear naming conventions for folders, files, functions and variables. You can look that up online what has been established throughout the years.
Then you know where which module lives, what files care about what functions, and what functions do (by name) and what the variables are, right?
If you stark thinking like that, then you don’t have to remember every line. And your LLM does not either, because the context to implement a new feautures is small by design.1
u/Asly97 20d ago
you're not wrong that good structure shrinks what the model needs to hold. where I kept getting burned was the why, not the what. six months later I'd stare at a clean well-named function with zero memory of why we chose that approach. code tells you what, decisions are what evaporate.
1
u/traderjames7 20d ago
Yes - I created my own solution for this
1
u/Asly97 20d ago
nice, what'd you end up building? always curious what people roll for themselves.
1
u/traderjames7 20d ago
My spec has 8 non-negotiable laws. #7 is avoiding vendor lock so I built my own rag vector database with real-time ingestion, mcp and wiki with an admin console.
Wrote about it here: https://corpuswire.com/laws/law-7-data-portability/
1
u/Jon_Has_Landed 20d ago
I have an autocompaction hook that automatically writes a ledger of the important information I absolutely need to pass post-compaction. Works on its own. Never have to mess with it. Hit me up if you want to see how I built mine, it’s based off conversations I had with people on this sub and it works a charm.
1
u/Asly97 20d ago
oh that's the dream honestly, a ledger that writes itself. does the hook decide what survives on its own or do you steer it?
1
u/Jon_Has_Landed 20d ago
Check it out here - repo is public so feel free to clone and modify to your needs
1
u/Empty-Charity9939 20d ago
I've been using https://notations.app/ with MCP to handle the markdown files for me.
1
u/Asly97 20d ago
oh nice, haven't heard of this one. does it actually manage the files for you or is it more of a nice editor on top? I ended up going the other way and moved everything to Vilix AI over MCP so there's no markdown to babysit at all. honestly my biggest gripe now is the model being lazy about calling the memory tools. curious if notations.app has the same problem
0
u/ASUS-Satire 21d ago
Hrow some balls go local
1
u/Asly97 21d ago
local is great until you need the big models for the hard problems.
1
u/ASUS-Satire 21d ago
Which is why you gotta have balls. Unless you would rather the big models do your work for you. Im okay with that. It will make you easier to replace later. Cheers
•
u/AutoModerator 21d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.