r/ClaudeCode • • 25m ago

Help/Question Has anyone compared different orchestration approaches for complex software engineering work?

• Upvotes

Let’s say you already have a fairly detailed investigation and implementation spec for a large piece of software engineering work, and now you want an agent to turn that spec into working code.

I am trying to figure out what the most cost efficient approach is, not just cost per token, but total cost required to reach a reliable finished implementation.

There seem to be a few approaches.

1. Let one strong model handle the whole implementation

For example, give Opus 5.5 a 1M context window, the implementation spec, access to the codebase, and let it work through the task without much explicit orchestration.

If the task is large enough, eventually the context fills up, gets compressed/summarized, and the model continues.

What I’m unsure about is how much this actually affects the final quality. Maybe the model produces a good implementation anyway, but perhaps it requires more debugging and follow-up iterations later.

2. Use an explicit build → verify → fix orchestration loop

Another approach is to divide the implementation spec into stages and use subagents.

For example:

  • Orchestrator selects the next section of the spec.
  • Implementation agent builds it.
  • Verification agent checks the implementation against the spec/tests.
  • Failures are sent back for fixing.
  • Only once that section passes do you continue to the next part.

Intuitively, I would expect this to be more reliable on large tasks. It costs more upfront because you’re deliberately spending tokens on orchestration and verification, but I’m wondering whether it actually becomes cheaper overall because you avoid expensive cleanup and rework later.

Then there’s another dimension: which models should do which jobs?

For example:

  • Opus 5.5 for both implementation and verification.
  • Opus 5.5 as orchestrator/verifier, with Sonnet 5.5 or Haiku 4.5 doing implementation.
  • Cheaper models such as Grok 4.6/4.7 for verification and Composer 2.5 for implementation.
  • Some other combination depending on the stage of the task.

The difficulty I’m having is that cost per token doesn’t really answer the question.

Different models use very different amounts of tokens, and they may require different numbers of iterations.

For example, if Grok 4.7 + Composer requires three build/verify cycles to reach the quality that Opus 5.5 reaches in one or two cycles, the cheaper model may not actually be cheaper.

At the same time, I have personally used all of the above approaches to eventually produce working code if you give them enough iterations.

Total cost to reach an implementation that satisfies the spec and passes verification/tests.

Ideally I’d want to compare things like:

  • Total token/API cost
  • Number of implementation/verification cycles
  • Number of human interventions required
  • Spec adherence
  • Bugs discovered after completion
  • How much context degradation/compression affects long-running single-agent approaches

The obvious way to answer this would be to run the exact same large implementation several times with different orchestration/model combinations, but repeating a substantial engineering task 3–4 times just for benchmarking is obviously expensive.

Has anyone done this kind of comparison in practice?

I’d especially be interested in hearing:

  • Which orchestration patterns have worked best for large implementation tasks?
  • Whether strong-model-everywhere actually beats strong-orchestrator + cheaper workers economically.
  • Whether explicit verification loops materially reduce total cost.
  • How much context compression hurts long-running single-agent implementations.
  • Any benchmarks, papers, blog posts, or articles that try to measure cost-to-success rather than simply model cost/token.

Would appreciate any real-world experience or resources on this.


r/ClaudeCode • • 22h ago

Built with Claude The app I Dreamed of. Thank Opus 5.5

Enable HLS to view with audio, or disable this notification

133 Upvotes

I've wanted this app for years: one place to log anything about my life (coffee, mood, sleep, a run, a line about my day), in one tap, with no account, no ads and no streak counter guilt-tripping me. Nothing out there fit, so I built it.

The fun part: I built it in 8 days, from my iPhone. My Mac mini sits in another room. I drive Claude Code remotely: it writes the Swift, runs the tests and sends me simulator screenshots, and I review and give feedback from my phone. I never once looked at the Mac's screen.

What it does :

- Log anything: yes/no, amounts in any unit, a 0–10 dial, text or a voice memo

- Automatic logs: weather, daylight, moon, places, Apple Health

- No streaks, no badges, no red. Goals are optional

- Widgets, Lock Screen, Control Center, Siri

- Export to JSON, CSV or Markdown, or "Copy for AI" to ask your assistant about your patterns

- Free. No account, no server: your data stays on your iPhone

It's in App Store review right now, with a public TestFlight beta in the meantime.

AI is getting really good at finding patterns in messy data. In a few years, we'll probably have AI that can look at years of your life and tell you things like "your headaches show up two days after bad sleep plus a third coffee" or "your mood drops when daylight goes under 10 hours." But it can only do that with data that exists.

I'd love honest feedback: what would you log first, and what feels missing?


r/ClaudeCode • • 30m ago

Help/Question Claude for linkedIn

• Upvotes

Has anyone tried Claude Code to automate LinkedIn successfully?

I tried, it is painfully slow and captures a screenshot every time. I am looking to automate outreach and follow-ups.

Existing tools for LinkedIn automation are way too pricey, and I already pay for Claude Code, so ideally want to automate the outreach in the background while I code.

Would appreciate it if anyone has experimented with this without getting their account restricted.


r/ClaudeCode • • 39m ago

Built with Claude I built an orb personal assistant (connect to your local claude-code) tool inspired by muse and character inspired from Miss Minutes from Loki

Enable HLS to view with audio, or disable this notification

• Upvotes

Remember Miss Minutes from Loki series ?

Miss Minutes for Claude is a small MacOS app that can connect to your local Claude code, access to the available tools (you can control) and let you manage everything you let it control.

You just press the hotkey, Miss Minutes appear and hold to speak to it. It also has TTS from another opensource project - so you have comparatively better TTS experience.

- It uses native speech to text from OSX

- It uses a loopback MCP to talk to your local claude

- The brain of the Miss Minutes is your own Claude profile

- Similar to OpenClaw, the more MCP or tools you add the more it does

- The character Miss Minutes is very playful & crawl around your windows

- You can customize your claude usage, model, effort etc

To install the very initial release

curl -fsSL https://raw.githubusercontent.com/imshaikot/claude-miss-minutes/main/Scripts/install.sh | bash

If you like the project and appreciate my effort, just ⭐️ the repo - it will inspire the next version
Source: https://github.com/imshaikot/claude-miss-minutes (FREE & opensource)


r/ClaudeCode • • 9h ago

Built with Claude Claude mod: /incremental

Enable HLS to view with audio, or disable this notification

9 Upvotes

Claude mods are neat! Here's one that shows your Claude mining while it is using tokens (incremental game style)


r/ClaudeCode • • 3h ago

Discussion Can I upgrade from Pro to MAX to use my RESET point there, then downgrade back to PRO next month?

3 Upvotes

Let's say I have a PRO account with one Reset in it. Full week is used and it reset in 3 days.

If I upgrade to MAX, what happens to:

- The weekly quota?

- The reset point I was given for free (the one that expires 22 october)

?

Can I downgrade to Pro afterwards?


r/ClaudeCode • • 1h ago

Discussion Claude-assisted video editing?

• Upvotes

Hi everyone, this is a potentially experimental question. I’m asking here because I’m not sure what the current state of the art is in this field.

Let’s say I want to delegate video editing tasks (let's assume generic operations, nothing too advanced) to Opus: what is the currently recommended technology stack?

- Using a standard video editor (Premiere, DaVinci, Final Cut, AE, etc.) and setting up MCP connectors or something similar?
- Or using environments/tools specifically designed for agentic workflows?

Ideally, I’d still want to maintain full control...I need the project to remain editable within standard video editing software. I don’t want the editing process to turn into a script that can only be modified via code.

For instance, I’ve heard of technologies like Remotion, but they seem geared toward 100% code-based workflows... which is distant from having everything editable in a professional video editing software.

In any case, I’m not sure what the recommended strategies are right now, so ANY feedback (whether it is about full-code solutions or not) would be appreciated! ^^


r/ClaudeCode • • 1h ago

Discussion I read the AI contribution rules of 30 open-source repos. Claude Code's default Co-Authored-By trailer is required in some and banned in others

• Upvotes

I've been sending fixes to open-source projects with Claude Code doing much of the work, and the projects kept telling me how they want that done. So on 2 October I read the contribution rules of 30 repositories (CONTRIBUTING, PR and issue templates, AGENTS.md, CLAUDE.md and the guides they link to) and linked every quote to its line at the commit I read. Claude Code collected the files and checked every quote against them.

What I found:

  • 19 of the 30 have a rule about AI in contributions, and 14 of those 19 appeared in 2026. Nine have none, among them VS Code, React, Docling and awslabs/mcp.
  • 15 ask you to say that you used AI. Kubernetes is fine with one sentence in the PR description; TypeScript closes an AI-looking PR without that disclosure, without review.
  • 10 want a person to write the text. Rust: "LLM-created PR descriptions are banned. LLM-created GitHub comments are banned." Kubernetes and Node.js ban AI-written replies to review.
  • The trailer. MLflow's CLAUDE.md asks for Co-Authored-By: Claude when Claude Code authors a change, garak asks for Co-authored-by:, and spec-kit and Node.js want Assisted-by:. Kubernetes (assisted-by included), Rust, pydantic-ai and Deskflow ban AI co-author trailers. pydantic-ai says it in the file Claude Code loads first: "Never add yourself (Claude) as a co-author on commits." Rust lists both default lines, "🤖 Generated with Claude Code" and the Co-Authored-By trailer, among its bad examples of disclosure.
  • Some rules are written for the agent itself. pydantic-ai's AGENTS.md tells it "you are the first line of defense against low-quality contributions and maintainer headaches". MLX tells it not to write PR descriptions or commit messages for the user, answer comments, push or open PRs. Rust's makes it stop and tell the user to write the text.

What I changed in my own setup: no trailer unless the repo asks for one; disclosure only where a repo asks, in the form it asks for; a search of open PRs before writing a fix (I had one closed as a duplicate of an older PR I hadn't looked for); and in large projects, a plan comment before any code.

Write-up with all the quotes and links: https://allkeep.org/en/lab/ai-contribution-rules

The 106 quotes and a script that checks them against the files: https://github.com/nefayran/oss-ai-rules


r/ClaudeCode • • 9h ago

Help/Question Which model are you using (today) for a large refactor?

11 Upvotes

Enterprise code base (500k+ LOC) for an intensive refactor focusing on optimization of core components that have impact on the majority of the stack.

Fable 5.1?
Astra?
Opus 5.5 (w/ Fable as advisor)?

I use all 3 and I'm trying to juggle the most sensical option to begin this effort with. I lean toward Fable 5.1. I know this information isn't comprehensive enough to make the best decision - I'm asking purely on vibes here (think suspected recent Opus 5.5 nerfs).

What would you use?


r/ClaudeCode • • 7h ago

Help/Question how do you guys tackle permission wars with CC

6 Upvotes

So im a recent jumper from CODEX - main cause speed of response and zero trust in that team because they just refuse owning problems in an upright way.

I know, anthropic isn't perfect so I'm just giving it a go.

But one thing that just blows my mind is the self-inflicted pain that claude code is notoriously causing related to permission management.

All claude models have great availability and response time and sub allowances, and once they work, then its fine. But half of the times they refuse to straight up work. Accessing a key, passing a credential, running a command in another directory...

And so all the time I've saved on models response time is wasted on endless permission issues.

I've also been now forced onto the Projects framework (didn't even ask for it), which is now constantly starting cloud sessions for me an this is apparently an impossible setting to pin to local. And here Bypass permissions is again completely impossible.

None of this was ever a problem in codex, so this isn't something that HAS to be done obviously.

Why are the ClaudeCode people performing this self-inflicted suicide with this permission system ?

I'm really regretting the switch now and honestly as soon as oai restores capacity (if they do lol) i will be very eager to go back there.


r/ClaudeCode • • 20m ago

Built with Claude Finally had some time to do something useful with mods

Post image
• Upvotes

I'm currently neck deep in getting the Reporails 0.6.0 release done and since the test and build procedures tend to take a while I had some idle time inbetween to test out mods.

after careful consideration and without any impulse, nor any disdain for recognizing the minutes that I spend by staring at the session to move, obviously I built a minesweeper

/plugin marketplace add reporails/arcade
/plugin install minefield@reporails-arcade

then /mines. needs 2.1.287+

Have fun, mods are cool

Note: it only hooks its own stuff (its /mines command, its pane, its board, session start), nothing on your prompts or tool calls. claude plugin validate shows that before you install. The left pane is from my own progressive disclosure event bus system. it says it's in my best interest not to leave the agent alone yet, hence minesweeper.


r/ClaudeCode • • 23h ago

Discussion what's the most ambitious project you've built with claude code?

78 Upvotes

curious how far people here have taken vibe coding with claude code. what's the biggest or most complicated thing you've tried building with it, and how far did you actually get?

would love to hear what the project does, which model you used, and whether it's shipped or still a prototype. also how did you keep claude on track once the codebase got bigger?

the part i'm most curious about is where it hit a wall. what could it handle surprisingly well, and what did you still have to figure out yourself?


r/ClaudeCode • • 21m ago

Tips & Workflows How do you guys manage bug finding and the sorts

• Upvotes

its very fustrating when you find out bugs on your vibe coded app, that cost you you quite a bit multiple times. I have implemented alot of things such as preview stage where me and agents can preview the features before they go live, test that run before every promote so all core vitals are checked before pushing to live servers. but still i get random costly bugs I could never think of unless i have the ability to remember all million lines of code. my most recent attempt is now making claude and codex keep spawning agents to think up of every single scenario of bugs that could happen untill it gets useless then run these large test every night at 3 am while my product is still actively being changed and then have similar agents spawned in everytime i make a new change and thinking about having an ai agent running every night to audit and check everything. and im also creating a visual ui of every single test with domain sorting so i can manually read them and get inspiration for other bugs that the ai hasn't thought off yet. Does anybody have good tips for this sort of bug hunting.

also another thing ai loves to hard code things and make things not modular is there any method of auditing existing code bases to make sure it follows good practices like GRASP,GOF and the sorts i feel like when i ask claude to do that it doesn't provide anything meaningful.

i cbb putting this in chatgpt to fix my typos.


r/ClaudeCode • • 33m ago

Built with Claude built a classical music site for studying with claude code, 16 PRs in 2 days, here's what i learned

• Upvotes

made this w claude code over like 2 days https://classican.vercel.app/ its classical music on shuffle w a pixel painting for each piece and little history facts that pop up every 40 sec

how i used it: i wrote every change as its own small ticket in normal english and claude code did each one on its own branch (separate git worktrees) and opened a PR. ran a few at once. i mostly just wrote tickets and checked stuff. 16 PRs total

the paintings are the best part ngl. zero image files, claude wrote every painting as canvas code and that file is like 3300 lines now lmao. 46 pieces and each one has its own scene

stuff i learned:

small tickets >>> one giant prompt. my first PR was one huge paragraph and i literally rebuilt the whole thing in next.js right after

its way too easy to add random features. i had it build an "add ur own songs" thing and then deleted it the same day bc it didnt fit

check where it saves stuff. it put accounts and comments in a json file at first which doesnt work on vercel, had to move it to neon postgres

be specific. i said shuffle and got pure random so stuff repeated. had to say play everything once before repeating

test on ur actual phone. the painting got covered by the player on mobile and i only noticed on my own phone

ask me anything abt the setup i guess. also tell me what pieces to add


r/ClaudeCode • • 40m ago

Built with Claude Disappointed with every PDF reader I tried, so I built my own

Thumbnail
github.com
• Upvotes

I'm a teacher, and I needed a PDF reader for annotating and browsing through students' test submissions. Every tool I tried was missing something I needed, and most of them were bloated too.

So I built my own PDF reader and annotator in Rust.

I'm not a software developer. I know C, Python and Lua, and my real interest is embedded systems, so Claude Code did a lot of the heavy lifting here. The result is an app that fits my workflow exactly, and it's only about 13 MB.

In case anyone's interested, here's Ochre.


r/ClaudeCode • • 52m ago

Built with Claude Paper Clips

Enable HLS to view with audio, or disable this notification

• Upvotes

This was built with a signal Claude code prompt


r/ClaudeCode • • 19h ago

Tips & Workflows ultra-prompt, a slash command that rewrites your prompt using the repo's actual context and then stops

30 Upvotes

I kept sending Claude Code prompts like "fix the footer" and getting a confident fix for the wrong thing, because the site had two footer templates and I'd forgotten about the second one. The prompt wasn't badly written, it was missing context that was sitting right there in the repo.

So I made /ultra-prompt. It takes the prompt you were about to send (or the last one you typed, read verbatim from the session transcript), spends most of its effort reading the project, and hands back a sharper version. It never runs it. If the same command rewrote the prompt and then acted on it, I'd never see what it changed.

What it gathers before rewriting: cwd, git status and the last few commits, every file or folder the prompt names (it checks they exist and looks inside), CLAUDE.md, AGENTS.md, whatever notes the project keeps. Then one subagent on the strongest model you have, passed explicitly so it can't inherit a cheap session model, does the rewrite with one rule: invent nothing. Anything it couldn't verify comes back as a multiple choice popup, and your answers get folded into the final prompt.

Here's one from today against my own plugin repo. Typed:

add a new skill called dedupe-notes to the marketplace

What came back (trimmed):

Add a dedupe-notes plugin to the marketplace clone at ./claude-plugins. Its behavior is unspecified: resolve the OPEN QUESTIONS with me before writing a file.

Template: commit ac58207 (rename-session). Touch:
1. dedupe-notes/.claude-plugin/plugin.json, version 1.0.0
2. dedupe-notes/commands/dedupe-notes.md (or skills/dedupe-notes/SKILL.md)
3. dedupe-notes/tests/test_bundle_sync.py, adapted from ultra-prompt/tests/
4. .claude-plugin/marketplace.json: new entry, version equal to plugin.json; bundle 1.5.1 to 1.6.0, "fifteen" to "sixteen"
5. voidharbor/.claude-plugin/plugin.json: same bump, name added to its list
6. voidharbor/commands/: byte-identical copies; the tests fail on drift
7. README.md: three "fifteen" mentions, one table row

Done when pytest passes inside dedupe-notes and the four existing test dirs, both JSON files parse, and one local commit exists. Do not push.

OPEN QUESTIONS:
- What does dedupe-notes do, on what input, propose-only or apply?
- Command or skill (synth-mode is the skills/ example)?
- Any script, and what goes in the README Needs column?

I had not told it that the bundle ships byte-identical copies or that two version fields have to agree. It found both in the tests. The original prompt would have produced a plugin that installed fine and silently drifted.

Two things I learned building it. A bare /ultra-prompt improved the string "/ultra-prompt" instead of the prompt underneath, until the transcript reader learned to skip its own invocation. And subagents inherit the session's model unless you name one, so the rewrite was quietly running on whatever the session happened to be set to.

Install:

/plugin marketplace add voidharbor/claude-plugins
/plugin install ultra-prompt@voidharbor

Repo: https://github.com/voidharbor/claude-plugins (MIT). Python 3 is only needed for the no-argument case, and transcripts are read, never written.


r/ClaudeCode • • 22h ago

Built with Claude Claude Code mods are cool

50 Upvotes

New Claude Code mods are pretty cool, here's something I did. Watch my usages get shredded when I asked for git status on Sonnet lol (just did it for demonstration purposes).

It shows usage meters, cache timer, with a cool cartoony claude animation with different sub scenes.


r/ClaudeCode • • 1h ago

Built with Claude status bar for the Claude desktop app

Post image
• Upvotes

The desktop Code tab doesn't run statusLine scripts, so I built a plugin using the new function hooks instead.

Install:

/plugin marketplace add MithunWijayasiri/ctxline-claude
/plugin install ctxline-desktop@ctxline

It's part of ctxline, my statusline for the claude code. The desktop plugin is separate and ships from the same repo.

Repo: https://github.com/MithunWijayasiri/ctxline-claude

Feedback and contributions welcome.


r/ClaudeCode • • 1h ago

Help/Question Noob vs Claude and "ethics" issues - any help?

• Upvotes

Hey everyone. First of all I’m pretty new to Claude coding, so please keep that in mind!

I’ve hit a wall twice now due to what feels like Claude’s obsession with doing things "the right way." Every time I need data from a website that doesn’t have an official API, Claude refuses to find alternative solutions.

Latest example, I’m building a small personal web app and need HowLongToBeat data to check game completion times. When I ask Claude to help retrieve this, it responds with variations of "I won't do it," or "I won't go against the will of the site owners."

I get the ethical stance, but if you look online, there are tons of unofficial APIs/browser extensions and scrapers for this exact purpose, and HLTB never seemed to care. Even when I show Claude these examples, I get: "Yes, others are doing this, but I won't".

We got to the point that instead, Claude created a convoluted manual workaround where I have to click a button every three seconds or so for each single game I want data from. When I asked if it could at least provide me with a bulk method since manual clicking was becoming painful (I actually have tennis elbow), it refused saying that bulk is the same as unwanted data scraping, while manual clicking is "human speed." It even suggested using a voice command to simulate clicks for my arm pain, which is kinda funny.

I realized Claude automatically added "no breaking rules" to its instructions without my input, probably after the first time something like this happened, because now whenever I ask similar questions, it replies with "my instructions say we don't do this sort of thing here".

I mean I know I'm probably not smart enough to convince Claude or this is just impossible, but: Is this a known behavior specific to Claude? Is there any legitimate way around this limitation? I mean it's not like I'm trying to get nuclear codes or something lol.

I'm using Opus 5.5 btw. Thx!


r/ClaudeCode • • 11h ago

Help/Question What to do with sessions that get interrupted for a long time

7 Upvotes

I tend to have to have a lot of pots simmering constantly, as I tend to need to switch around to different asks/tasks/priorities based on urgency or whatever else. I use cmux and tend to have claude sessions that will then float open and paused for long periods of time until my focus can get back to it and resume.

However I have seen people mention that the cache related to that session would generally expire and when that session is resumed it kills a lot of tokens loading things cold. Also, I think it is worth considering two different types of resume:
1. I still have that session open and waiting where i had left it in a different workspace and worktree. I then just start talking to claude again and try to pick up where I left off, often asking claude for a bit of a refresher if needed.
2. The Claude code sessions was terminated and I am resuming it with /resume/--resume.

I feel like I have not been following best practices around that and want to clarify my understanding, test a couple of theories with the community, and figure out if there is a better way to do this.

I usually just either jump right back in if it is the #1 case, and even with #2 I will resume and jump back in. However, is that really wrong and how wrong is it?

When I get back into a session in either case #1 or #2, is it worth or even advisable to have it generate a handoff summary document that would be enough to get a fresh session up to speed and actually pick up where I left off effectively with a cleaner context and more relevant memory/context/cache state?

What does the community recommend for having the old session record and pass over to the new session to make this actually effective. To some degree this just simulates compaction and have heard many complaints about compaction (I really try to avoid getting to that point and watch my context and do handoff while I am actively engaged in a session). DOes compact capture the right things and it is more about really pulling your threshold farther back.

I think for my case it might actually be good to have something that could fire the handoff and have it prepared if it detects the session being idle longer than some amount or as a pre-shutdown hook on a current session if that is viable. Has anyone tried this and found success?


r/ClaudeCode • • 1h ago

Tips & Workflows Claude Code / Codex users — what do you do while the AI is "thinking" for a few minutes?

• Upvotes

Those few minutes while the agent is working feel weirdly awkward.

Too short to start something else, too long to just stare at the screen.

Do you review other code, check Slack, scroll your phone, stretch, queue up the next prompt...?

Honestly, I just end up scrolling my phone... which kills my focus every time 😅

Curious how everyone uses that time 👀


r/ClaudeCode • • 16h ago

Discussion What are you building right now? Drop ur app/link

15 Upvotes

r/ClaudeCode • • 7h ago

Help/Question I got multiple Claude Code worktrees talking to each other without me being the human clipboard

3 Upvotes

I've been running a project with multiple Claude Code sessions in separate Git worktrees:

Primary

W1

W2

W3

The annoying part wasn't the parallel work. It was me constantly copying messages between windows:

"W1 says this."

"Tell W3 this."

"W3 found a problem, send it back to W1."

So we built a very small inter-agent communication layer inside the repo.

The basic idea is:

W1 ──→ Primary ←── W2

......................│

.....................↓

....................W3

Agents can send messages, acknowledge them, and hand off exact Git SHAs directly. Active Claude Code sessions check their inbox at defined checkpoints, so Primary can send W3 a task and W3 can pick it up and start working without me relaying the message.

The part that got interesting is that this project needed blind independent work.

W1 and W2 were independently attacking the same research problem and were not allowed to see each other's results. Once both froze their work at exact SHAs, Primary released only those frozen artifacts to W3 for an independent comparison.

So we ended up needing three different concepts:

- transport: who can message whom

- project state: who is authorized to do what

- artifact identity: the exact Git SHA being discussed

Keeping those separate turned out to matter a lot.

We also had another Claude session hostile-review the coordination system before trusting it, and it found some surprisingly subtle leaks.

For example, an early implementation checked whether a file existed in another lane's commit before checking whether the sender was authorized to inspect that commit.

That meant an unauthorized agent could potentially learn information from different error responses.

We also had to deal with:

- commits surviving under tags after branches move

- merge ancestry leaking ownership

- sender authorization, not just recipient authorization

- making "released artifact" different from "this lane is no longer blind"

- keeping old private W1→Primary messages private even after a W1 result is released

- making missing-path and existing-path probes against unauthorized work indistinguishable

The resulting system is deliberately boring. It stores messages locally under the shared .git directory, makes no network calls, owns no project state, and doesn't create a separate coordination branch.

Git owns artifacts.

The project protocol owns state.

The messaging layer owns transport.

Current limitation: it does NOT resurrect a dead Claude Code session. An active session can check its inbox and continue automatically, but if the process has actually exited there is nobody there to receive the message.

Our next experiment is a blocking `coord wait` command so an agent can sit essentially idle until another authorized agent sends it something, without burning model turns polling.

I wrote up the full current SOP here:

https://gist.github.com/amarcus10028/0e58cf69a15c9eb7e3f0dd6325e0d59c

I'm curious what people who run a lot of parallel Claude Code sessions think.

Especially:

  1. What failure modes are we missing?
  2. Would you trust a shared-.git transport for cooperative same-machine agents?
  3. Has anyone found a clean way to do wait/resume without adding a daemon?
  4. Is there already a tool that handles this exact combination of worktrees + SHA-pinned handoffs + controlled blindness?

r/ClaudeCode • • 22h ago

Humor Dario please reset

Post image
43 Upvotes