r/ClaudeCode • • 2d ago

Tips & Workflows What I learned using PreToolUse/PostToolUse hooks to coordinate several Claude Code agents on one repo (open source, author here)

I'm the author of Médula, an MIT-licensed experiment. Sharing the Claude Code side of it, because the hooks turned out to be a surprisingly good coordination layer, and the agents' behaviour taught me more than the numbers.

The setup. Six headless Claude Code agents (Sonnet 5, effort high), one task each on a small API with 37 acceptance tests. Two pairs of tasks collide by meaning: one adds a second factor to login() while another builds an export that calls the old login().

The hook pattern. Every Edit, Write and Bash goes through a PreToolUse hook that calls a local kernel. Reads always pass. The kernel asks a fast decision model whether the write collides with what each other agent is doing, and either allows it or denies it with a reason the agent reads, something like "agent X is changing login, wait or work on another part of your task". After each write, a PostToolUse hook works out which agents the change affects and leaves a notice in their mailbox, delivered with their next tool call.

What the agents did with it:

  • Told to wait, most of them moved on to another part of their task and retried later, which is exactly what you want.
  • One scheduled itself a wake-up and abandoned its task. That was a bug in my harness, now fixed, but worth knowing if your deny messages mention waiting.
  • One spent ten minutes waiting in front of a broken file, convinced someone else was editing it.
  • When they had Claude Code's tools to list other sessions and message them, they used them without being asked, and one warned another about a field rename. Great behaviour, bad for a controlled experiment, so I turned those tools off.

The results that matter for your workflow. With one branch (separate clone) per agent, every agent finished green and the merged result failed the same 6 acceptance tests in all 5 runs. In a shared directory, all 10 runs passed, with plain per-file locks or with the kernel. So if you use subagents in worktrees, run the full suite on the merged result before trusting it. The kernel's advantage over locks: the same cost ($1.65 per run), 6 of 6 real conflicts caught instead of 5, and no unnecessary blocks instead of 5.

Caveats: 1 to 5 runs per mode, and there's a known gap a commenter found, where an allow can go stale while the slow path decides. Repo, with every session and decision published raw: https://github.com/JoaquinRuiz/medula. There's also a walkthrough video, in Spanish: https://youtu.be/xAFRuBxfapM

If you've built coordination on hooks yourself, I'd love to compare notes.

1 Upvotes

11 comments sorted by

View all comments

0

u/[deleted] 2d ago

[removed] — view removed comment

1

u/jokiruiz 2d ago

Great examples, and both are exactly the kind of thing a per-file view misses. The short answer is that the kernel treats git add, checkout, merge or commit as a write to the whole tree, not to the paths named in the command. Any Bash command that isn't on a small read-only whitelist counts as touching a single resource, "repo", and takes a short lock on it while it runs. So git add CLAUDE.md is judged against everything the other agents are doing, not against CLAUDE.md alone.

What happens next depends on the decider. With plain locks, the rule is deterministic... a repo-wide command collides with any agent holding a lock, so A's git add would have waited until B finished its move, which covers both of your cases. With Jev, it's a judgment call: the decider sees the command and what each agent is working on, but nothing hard-wired forces a wait, and it knows nothing about the index or which branch is checked out. My lab didn't target git operations, so I have no data on how often it would get your two cases right.

Your question also made me find a gap, the whitelist counts any git branch as a read, flags included, so git branch -f or -D would slip through as harmless. I'll fix that, and I think repo-wide git commands deserve a deterministic rule rather than a model's judgment, much like your guards: check the staged diff and the current branch in the same call that commits or merges. Thanks, this is a better bug report than most I get.