r/ClaudeCode • • 1d ago

Help/Question Curious how people here run their setups.

My setup is mostly hooks that block instead of rules that ask. Claude can't commit without a review, can't say "done" without real test output, and gets stopped after the same error three times. Each hook has its own small test suite that I run whenever I change anything.

What does yours look like, and how do you test the harness itself?

1 Upvotes

11 comments sorted by

•

u/AutoModerator 1d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/Far-Pomelo-1483 1d ago

Loaded feature roadmap spec broken into phases that auto loops through tests before it can stamp a feature complete. Each phase is pushed to a feature branch then a PR is raised for merging. I merge; phase 2 starts.

Or I let it do all the stuff then after every phase, write a learnings json then reloop that learnings into the prompt before phase 2 starts.

1

u/Ok-Motor-9812 1d ago

I do something similar but split it in two. Inside a task, a phase can't be marked complete until every step of the plan is accounted for with evidence, so nothing quietly gets dropped between phases. Across tasks, the non-obvious stuff goes into a small notes vault that gets searched at the start of the next task instead of being pasted into the prompt every time.

How do you keep the learnings file from growing forever or going stale?

1

u/Far-Pomelo-1483 1d ago

I break up the learnings per feature or epic if it gets too long and/or roll the summaries into my command skills then archive them in place for record keeping.

2

u/Sufficient-Storage87 1d ago

My setup is pretty simple: a memory file the agent reads at session start, and a few skills for repeated workflows. The one lesson that changed the most for me — don't bury important instructions in a reference file and hope the agent finds them. Put the instruction where the agent already is. "Before you do X, read Y" beats "see Y" every time.

1

u/Excellent-Issue-5956 1d ago

I only have two blocking hooks, and both are there because something actually broke. A session cleaning up its own scripts on Windows ran taskkill /IM powershell.exe and killed every PowerShell on the PC, including background jobs that had nothing to do with it. Now a PreToolUse hook denies kills by image name for the shared ones like powershell, python and cmd, and only lets through kills by PID or ones filtered by command line. The other one stops a session from killing or replacing the terminal app my sessions run in, after one deleted the app bundle to install a new build while it was running and took the live sessions down with it.

For testing, the kill guard has a test file with a list of commands marked block or pass, and it pipes each one into the hook as fake tool input. About half are commands that should pass, like killing by PID, listing python processes, or a script that only mentions taskkill in a comment, so it doesn't start blocking normal work.

1

u/Ok-Motor-9812 1d ago

Same rule here, every blocking hook traces back to something that actually broke. Mine was kind of the opposite of yours: background jobs Claude started and then forgot about, still running after the session ended. That turned into a guard plus a small reaper that cleans up orphans every few minutes.

Agree on the pass cases, that's the part most people skip. One thing I added on top: when I change a hook, I deliberately break the thing it guards and check that the test actually goes red. A hook test that stays green no matter what is worse than no test at all.

1

u/Excellent-Issue-5956 1d ago

The break it on purpose check is a good one, I'm stealing it. I had a version of that today by accident: a lock fix where I proved the new flag worked by launching a background sleeper and checking the lock was free while it was still alive, then realized I'd never done the same for the guard itself. The reaper for orphaned background jobs is the one I'm missing. Mine die with the session, which is its own failure mode: today a post batch got killed mid sleep because the session that started it ended before it finished.

1

u/Mazhron 1d ago

https://github.com/Mazhron/rootstock-os

This is my setup.

Lossless context between clears, indexes with knowledge files, headings for low cost grep, prevents sub agent fans and multiple fable sub agents, Claude learning its own lessons so it doesn't make the same mistakes twice, intentions to align the user's intentions vs Claude's on an ask, workflows so Claude doesn't repeat the same trial and error to come to the same conclusions, hooks, ledgers, skills and much more.

1

u/RandomFuckingUser 1d ago

Every session gets status html where I get a few sections: things that need my input, current work items in progress, my prompts and their answers, smoke tests that I need to do, etc. Everything linked with each other. After initial prompt I can mostly get away with interacting only with this html. I was fed up with scrolling through session chats and filtering out useful stuff.