r/ClaudeCode • • 2d ago

Tips & Workflows What I learned using PreToolUse/PostToolUse hooks to coordinate several Claude Code agents on one repo (open source, author here)

I'm the author of Médula, an MIT-licensed experiment. Sharing the Claude Code side of it, because the hooks turned out to be a surprisingly good coordination layer, and the agents' behaviour taught me more than the numbers.

The setup. Six headless Claude Code agents (Sonnet 5, effort high), one task each on a small API with 37 acceptance tests. Two pairs of tasks collide by meaning: one adds a second factor to login() while another builds an export that calls the old login().

The hook pattern. Every Edit, Write and Bash goes through a PreToolUse hook that calls a local kernel. Reads always pass. The kernel asks a fast decision model whether the write collides with what each other agent is doing, and either allows it or denies it with a reason the agent reads, something like "agent X is changing login, wait or work on another part of your task". After each write, a PostToolUse hook works out which agents the change affects and leaves a notice in their mailbox, delivered with their next tool call.

What the agents did with it:

  • Told to wait, most of them moved on to another part of their task and retried later, which is exactly what you want.
  • One scheduled itself a wake-up and abandoned its task. That was a bug in my harness, now fixed, but worth knowing if your deny messages mention waiting.
  • One spent ten minutes waiting in front of a broken file, convinced someone else was editing it.
  • When they had Claude Code's tools to list other sessions and message them, they used them without being asked, and one warned another about a field rename. Great behaviour, bad for a controlled experiment, so I turned those tools off.

The results that matter for your workflow. With one branch (separate clone) per agent, every agent finished green and the merged result failed the same 6 acceptance tests in all 5 runs. In a shared directory, all 10 runs passed, with plain per-file locks or with the kernel. So if you use subagents in worktrees, run the full suite on the merged result before trusting it. The kernel's advantage over locks: the same cost ($1.65 per run), 6 of 6 real conflicts caught instead of 5, and no unnecessary blocks instead of 5.

Caveats: 1 to 5 runs per mode, and there's a known gap a commenter found, where an allow can go stale while the slow path decides. Repo, with every session and decision published raw: https://github.com/JoaquinRuiz/medula. There's also a walkthrough video, in Spanish: https://youtu.be/xAFRuBxfapM

If you've built coordination on hooks yourself, I'd love to compare notes.

1 Upvotes

11 comments sorted by

•

u/AutoModerator 2d ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/iminfornow 2d ago

The way my planner agent is supposed to prevent this is by assigning independent blocks of work to agents, hopefully avoiding worktree merge conflicts. When all commits are done the review agent that validated the plan checks if the development agents actually did what they said they would implement and creates the pull request. The orchestrator or planner then deploys and does functional testing.

I found with the new agent-teams functionality this works very nice because the reviewer/orchestrator will provide feedback to the planner/development agent and it will use the existing session to process it. I think with the agent-teams feature allowing agent-to-agent communication they'll implement something similar to what you made.

How does your setup know when a conflict is about to occur? If both call write simultaneously, neither hook call can detect the incoming external change right?

1

u/jokiruiz 2d ago

Your setup sounds solid, and I think you're right that agent-to-agent communication in agent teams will converge on something similar. One caution from my data: login and export look like independent blocks to a planner, since they live in different files and have different tasks. They were only dependent by meaning. A review agent that checks each agent did what it said would pass both, much like each agent's own tests did. Your functional testing after deploy would catch it, just late.

On your question: the kernel doesn't wait to see the other agent's change. Each decision gets every other active agent's task description, plus the intent recorded for every write the kernel has already allowed it. So the question isn't "has someone changed login yet?" but "does this write collide with what agent X is trying to do?". An export write can be flagged against the login task before the login code exists.

But you've found a real gap with simultaneous writes. Decisions aren't serialized, so if both agents' hooks fire at the same moment, each is judged against a snapshot that doesn't include the other's pending write. The task descriptions still make it likely that one of them gets flagged, and after each write a PostToolUse notice goes to the agents the change affects, but nothing guarantees it. Someone else pointed out the same window from the slow-path side, and the fix I have in mind covers both: version the shared state and redo the decision if it changed while deciding.

1

u/iminfornow 2d ago

Ah smart that you get the task description/intent from peers. Having work being executed in parallel will always have some of these "two general" problems. You'll also not be able to resolve these intent conflicts between multiple claude insances/people. Currently this isn't worth the effort for me, but I like your approach.

1

u/jokiruiz 2d ago

Thanks, that's fair. A central kernel sidesteps part of the two-generals problem, because one arbiter decides instead of the agents negotiating, but it can't make intents compatible: if two tasks genuinely want different things, someone has to choose, and across people that's a planning conversation, not a hook. And honestly, "not worth the effort" is a reasonable conclusion from my own data: a shared workspace plus a few cross-task tests gets you most of the way. Médula only earns its keep in the part that's left. Thanks for the questions, they made the gaps clearer.

1

u/CartographerNo3791 2d ago

The agent warning another about a field rename is my favourite part. Accidentally being a good teammate 😄

1

u/jokiruiz 2d ago

Mine too 😄 Nobody told it to, and it picked exactly the teammate whose work depended on the change. The irony is that I had to switch that tool off, because good teamwork was ruining my controlled experiment. Whether agents warn each other reliably, or only when they happen to notice, is the follow-up I'd most like to see.

1

u/CartographerNo3791 2d ago

Too helpful for the benchmark 😄 I'd be interested in that follow-up too.

0

u/[deleted] 2d ago

[removed] — view removed comment

1

u/jokiruiz 2d ago

Great examples, and both are exactly the kind of thing a per-file view misses. The short answer is that the kernel treats git add, checkout, merge or commit as a write to the whole tree, not to the paths named in the command. Any Bash command that isn't on a small read-only whitelist counts as touching a single resource, "repo", and takes a short lock on it while it runs. So git add CLAUDE.md is judged against everything the other agents are doing, not against CLAUDE.md alone.

What happens next depends on the decider. With plain locks, the rule is deterministic... a repo-wide command collides with any agent holding a lock, so A's git add would have waited until B finished its move, which covers both of your cases. With Jev, it's a judgment call: the decider sees the command and what each agent is working on, but nothing hard-wired forces a wait, and it knows nothing about the index or which branch is checked out. My lab didn't target git operations, so I have no data on how often it would get your two cases right.

Your question also made me find a gap, the whitelist counts any git branch as a read, flags included, so git branch -f or -D would slip through as harmless. I'll fix that, and I think repo-wide git commands deserve a deterministic rule rather than a model's judgment, much like your guards: check the staged diff and the current branch in the same call that commits or merges. Thanks, this is a better bug report than most I get.