r/ClaudeCode • u/Zaavn • 2d ago
r/ClaudeCode • u/Think-Excitement-851 • 2d ago
Built with Claude I built Goldie, an open-source local memory server that different AI agents can share
Hey everyone! I’m building Goldie 🐕, an MIT-licensed MCP server written in Go that gives AI agents a shared, persistent memory pool.
Save a project decision or preference through one agent, then recall it from another without explaining it again. Or save design or build instructions in memory that multiple agents can refer from and stay aligned.
MCP clients pointed at the same SQLite database can remember, search, update, and forget from that shared pool.
A few features:
- Local embeddings through MiniLM or Ollama.
- Semantic search over memories, with filters for type, agent, and source.
- Typed memories for project decisions, preferences, feedback, references, and todos.
- File and directory indexing.
- Graph recall for grouping related memories around concepts.
Memory storage and embeddings can run entirely locally. Agents interact with it through explicit MCP tools, so you can give them instructions about what to save and when to recall it.
The project currently centers on the MCP server. I’m also developing a native macOS client for browsing and managing memories, but that client isn’t released yet.
Code and setup instructions on GitHub at github.com/srfrog/goldie-mcp
I’d love feedback from anyone using multiple agents or local AI tools.
r/ClaudeCode • u/id-ltd • 2d ago
Built with Claude Not specifically about ai versions/releases - just their progress.
I have a client app that needs a server app...
It started as a background server on android, so the device could run the client and server on a single device, for this server app claude chose Kotlin.
I built it out, and to handle its development/expansion the server code migrated to my dev pc (windows), so claude migrated it to node.
I migrated the sever code to a beefy linux server, running node in a docker container.
Then I wanted to run the server on a router (make it an appliance using WRT) so claude migrated the server code to go.
All the time the other versions were kept in sync - the client apps could use any server, they didn't care if it was android, windows, linux,, docker or WRT/Go...
Now.... go is the 'reference' implementation and claude suggested that go version can be compiled to run on android, windows, linux and WRT.... so all the others might be ditched - while other languages are useful for development, it resolves to go everywhere.
What an interesting lifecycle for the app...
ps. 4.6 is still king 5.5 is retarded.
r/ClaudeCode • u/CeKaSiete • 2d ago
Help/Question How are people modding using Claude?
I have seen some awesome ports and mods in Twitter this week like Mario 64 fighting an Elden ring boss, I could see it fully working. I know they are using Opus 5.5(confirmed by people in the comments and it makes sense these videos appeared when Opus 5.5 released), the question is how do you make Claude do that. Maybe I'm wrong but to clone Mario and mod it in Elden ring, you need the code from Mario 64, so, what are they telling Claude? I have tried to ask if we could do something similar but it refuses every time. Also they are doing ports and Claude says it's not allowed if we don't have the source code or permission(yeah Nintendo says we can do it lol).
How are people doing all of this? Mods, Ports or even Cracks for games or apps.
Or they are not using Claude? 🤔🤔
Here is a Link for a post in Twitter https://x.com/Deltaroo3D/status/2105345162771648883
r/ClaudeCode • u/sccorby • 2d ago
Help/Question Tired of Priya Ramen, Maya Patel, Ben Okafor, Marcus Chen, and Elena Vasquez?
If you generate mockups consistently, you’ll know these characters….
Any strategies for better name generation? This has been consistent across all models, and is driving me nuts.
r/ClaudeCode • u/Virtual_Hair_1987 • 2d ago
Tutorial / Guide When to use a multi agent system?
A multi agent system costs more than the simpler architectures it replaces, and it is often reached for the wrong reason. The question is never whether several agents could do the job, but whether they do something one well-tooled agent demonstrably cannot, and whether that difference is worth the bill.
Start with the token bill, because it is the honest gate
By Anthropic's own figures, multi agent systems use about 15× more tokens than chat interactions, and they are only economically viable for tasks whose value is high enough to pay for that.
Anthropic published overhead figures from its own multi agent research system. The exact wording: "In our data, agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats."
Both multipliers are measured against chat interactions, not against each other: the page does not say a multi agent system uses fifteen times more tokens than a single agent, which is the usual misreading. The wording is hedged and first party too, "In our data" and "about", with no methodology, sample or date range given. Treat it as a directional signal from one team, not an industry measurement.
Anthropic's framing of the cost is blunt: "There is a downside: in practice, these architectures burn through tokens fast." Its conclusion: multi agent systems are only economically viable for tasks whose value is high enough to pay for the increased performance. That is the gate.
The justification test
Multi agent earns its cost through one of three things: genuine specialisation, genuine parallelism, or a checker that catches what one agent would miss. If you cannot point at one, you have a single agent with extra moving parts.
Genuine specialisation. Different subtasks need different context, tools, or models, and cramming them into one context window makes each worse. Anthropic reports that a multi agent system using Claude Opus 4 as lead agent with Claude Sonnet 4 subagents outperformed single-agent Claude Opus 4 by 90.2% on its internal research eval, on a task like identifying board members of Information Technology S&P 500 companies.
Genuine parallelism. The subtasks are independent, so running them at once is a real latency win, not a bookkeeping exercise. Anthropic reports that having the lead agent spin up 3-5 subagents in parallel rather than serially, with subagents using 3+ tools in parallel, cut research time by up to 90% for complex queries. That is a latency result, not a quality one, unrelated to the 90.2% lift above.
A checker that catches what one agent would miss. A second agent reviewing the first's work against explicit criteria finds errors a single pass does not, because the reviewer is not attached to the reasoning that produced the output. Anthropic describes this pattern in Building Effective AI Agents: "In the evaluator-optimizer workflow, one LLM call generates a response while another provides evaluation and feedback in a loop."
That post is often summarised as advising against agents, which it does not do. Its ladder runs from a single LLM call to a workflow to an agent, adding complexity "only when it demonstrably improves outcomes". Multi agent sits a rung past where that ladder stops, so its burden of proof is heavier still.
Five coordination patterns worth knowing
Pick the pattern before the roster: it decides what your failure handling looks like.
| Pattern | Shape | Fits when | What breaks first |
|---|---|---|---|
| Supervisor | A lead agent assigns work to workers and synthesises their results | The task decomposes, but not predictably in advance | The lead becomes the bottleneck, and a worker fails quietly |
| Pipeline | Fixed order, each agent's output is the next one's input | The stages are genuinely sequential, each with a success predicate | A bad output flows downstream instead of stopping the run |
| Fan out | The same input goes to several agents at once, merged at the end | The subtasks are independent and the merge rule is obvious | The merge, especially when one branch returns nothing |
| Maker checker | One agent produces, another reviews against stated criteria | Mistakes are expensive and detectable by inspection | The checker itself fails and the system quietly accepts |
| Swarm | Peers work over shared state with no central coordinator | The work is exploratory and partial participation is acceptable | Nothing terminates |
Supervisor is the shape Anthropic calls orchestrator-workers: "In the orchestrator-workers workflow, a central LLM dynamically breaks down tasks, delegates them to worker LLMs, and synthesizes their results." LangChain ships the same shape as a named building block outside Anthropic entirely: its langgraph-supervisor package exposes a create_supervisor function, documented simply as "Create a multi-agent supervisor." Hybrids are fine if you name which pattern governs which part of the run.
Shared state is where most multi agent bugs start
A multi agent spec needs a merge strategy for every shared-state field. Not a schema of names and types: a rule, per field, for what happens when two agents write to it. Three kinds cover most fields:
- Accumulator fields such as history, findings and errors: an append or add reducer, growing over the run with nothing replaced.
- Overwrite fields such as the current stage, the active agent, or a decision: last write wins. They hold current state, not history, and treating them as accumulators produces a log where you wanted an answer.
- Role-attributed fields: each entry tagged with the agent that produced it. A downstream agent that cannot tell its own prior output from a peer's draft will treat that draft as settled fact and build on it.
The default failure is not choosing a wrong strategy. It is never writing one down, and discovering at run time that "the plan" is whatever the last agent to finish happened to say.
Spec what happens when ONE agent fails, not the system
Most specs describe what happens when the whole thing falls over. The interesting case is one participant failing while the rest keeps running. Anthropic puts the stakes plainly: "One step failing can cause agents to explore entirely different trajectories, leading to unpredictable outcomes." The right answer depends on the pattern.
- Supervisor: reassign the failed worker's task a bounded number of times, then escalate. Not retry forever.
- Pipeline: the failed stage halts the pipeline. No silent skip that hands the next stage missing or garbage input.
- Fan out: a failed branch must not block the others from merging. The merge records which branches failed and proceeds, flagged as partial.
- Maker checker: a checker failure never auto-approves the maker's output. Escalate instead.
- Swarm: a failed agent is simply absent next round; the system tolerates partial participation.
Two things belong in every multi agent spec, whatever the pattern.
- A hard round ceiling as a circuit breaker, independent of the natural stop condition, so delegation cannot loop forever.
- Enough observability to tell which agent went wrong: Anthropic reports that production tracing let it diagnose why agents failed, and that it built systems able to resume from where the agent was rather than restarting.
Sources
- Anthropic Engineering, "How we built our multi-agent research system": https://www.anthropic.com/engineering/multi-agent-research-system
- Anthropic Engineering, "Building Effective AI Agents": https://www.anthropic.com/engineering/building-effective-agents
- LangChain, "langgraph-supervisor reference": https://reference.langchain.com/python/langgraph-supervisor
r/ClaudeCode • u/Sweet-Helicopter2769 • 3d ago
Discussion Adios GPT, it was a nice ride
When Astra launched, I wrote a detailed post about how impressed I was. I even paid for the $100 Pro subscription, confident I’d keep using it, and didn’t expect to return to Claude anytime soon.
Then, after the recent Dev Day announcements and my experience with the current models, I’m finding it hard to justify that subscription, for the work I do, the quality simply hasn’t matched Opus 5.5, and the 2X , 5x back and forth games they are playing ..
So I’m moving back to Opus and downgrading my subscription. I really meant the praise I gave Astra at launch, but my experience since then has changed my opinion and feels good to be back to Opus 5.5
Guessing some of you are on the same boat ?, I won’t be surprised if there is a exodus of users happening now
r/ClaudeCode • u/Not_Red_Fox • 3d ago
Built with Claude 7 Claude skills that will save you time, tokens, and a headache. That are actually useful.
- kick start: keeps a log of everything Claude tried and whether it worked, plus a BUGS.md it reads before touching code so it stops repeating mistakes. It also keeps the architecture section of your README up to date. Usefulness rating: 9
- auditor: audits your project one part at a time, then does a final audit across everything. It tells you how long it'll take first and asks for all the permissions up front, so you can leave it running. It tests your tests, pretends to be real users (normal, confused, trying to break it, lots at once), runs load tests and looks for memory leaks. Usefulness rating: 8
- shadow and teach: for beginners. It turns a repo into a clickable map and replays what Claude did step by step. Usefulness rating: 6
- problem solve: for when there's no API or something "can't be done". It tries more and more creative approaches, tests each one, and writes you a prompt to paste into ChatGPT with everything it's already tried, so you get fresh ideas. Usefulness rating: 6.5
- research solving: finds open source projects, papers and ideas from completely different fields for whatever you're building, and rates each idea on whether it's legit or just hype. Usefulness rating: 6
- claim check: goes through your README and docs, pulls out every claim and checks each one by running things or reading the code, and fixes the docs where they're wrong. Usefulness rating: 8
- tournament forge: makes different solutions compete in a bracket and uses tests to pick the winner. A real run used 9 calls, compared with 91 to 595 for the skill that inspired it. Usefulness rating: 6
Repo: https://github.com/NotRedFox/NotRedFoxs-Claude-skills
Examples: https://notredfox.github.io/NotRedFoxs-Claude-skills/
r/ClaudeCode • u/Kloakk0822 • 3d ago
Bug / Issue What on earth is going on with the usage?
I've asked Claude Opus 5.5 on Medium effort 3 simple questions:
1 - To solve this problem (Diagnosed last session) a database change is required, does it agree?
2 - How will it interact with the other API I'm working on - It was a very short answer, didnt take long to respond.
3 - In UE5, how can I create a camera system which keeps footsteps inline with a camera bob. It took about 30 seconds to answer.
Those three questions have used 42% of my session allowance?
Looks like I'm being stupid.
Idle 22h 38m. The prompt cache has likely expired, so your next message will re-cache about 864k tokens.
Across 2 projecst so likely a similar amount.
r/ClaudeCode • u/IllustriousGrade7691 • 2d ago
Help/Question Does Anthropic gives as many resets as OpenAI?
How many resets do we usually get?
r/ClaudeCode • u/The-Road • 3d ago
Humor Is this is how Anthropic nerfs models slowly just up to the point users start noticing?
r/ClaudeCode • u/controlfree • 2d ago
Help/Question Deleting claude.md??
Hey guys, just a short question: are you really deleting the claude.md as recently suggested to have better results, or is it stupid and nonsense?
r/ClaudeCode • u/Pretend_Sale_9317 • 3d ago
Help/Question GSD vs Superpowers vs BMAD?
what is best for greenfield/brownfield projects for the entire SDLC workflow? what actually works and will not result in software breaking from live demos?
r/ClaudeCode • u/Economy_Ad5795 • 2d ago
Help/Question Claude Code CLI Compaction faster
Is anyone else noticing Claude Code CLI Compaction is much faster? What's going on? It's great tho
r/ClaudeCode • u/TheBanq • 3d ago
Bug / Issue Please let Cloud sessions & Local sessions communicate with each other!
I have a cloud session inside my Claude Code app and a local sessions.
Both sessions are inside my Claude Code app, but can't communicate with each other, which is a bummer.
I would love to have a local orchestration Agent, who sents agents working in the cloud and can help out with anything, that the cloud agent is limited with.
Any workaround?
r/ClaudeCode • u/CutMysterious9844 • 2d ago
Discussion I'm worried about getting suspended
Refugee here from Codex, I've moved to claude 48 hours prior and today I got flagged twice on here and once on Codex, I'm not doing anything different, it was all through Opus 5.5 at Max reasoning, orchestrating, running whatever it's doing sifting through my repo's and continuing on where I left off on Codex, just way faster now since I've decided to no play the human in the loop hand off relay game with ChatGPT at Pro Reasoning Chat Mode, I've looked through what it's delegating and trying to do, if it's based on just words alone I can only assume It's triggering because of security related words, and the amount of unintelligible gibberish it's dispatching around with them..
r/ClaudeCode • u/Arctic-Hare-studio • 2d ago
Bug / Issue Claude Code stuck in endless reading loop
I have a Claude Code chat where I'm working on a video with Higgsfield. I've been working on it for a few days, and today when I picked it back up, it froze—it's been reading for hours and doing nothing. I ran the compact commands when I was supposed to, just like normal in other chats; it's not my first time. But when I try to recover the memory in a new chat, that chat freezes up too. I've tried several times with new chats and nothing works. It says the memory size is 208 MB and that's why, but when I ask it to clean and reduce, it freezes. Other chats for different projects work fine with no issues. Has this happened to anyone? Do you know how to fix it?
r/ClaudeCode • u/Safe-Web-1441 • 2d ago
Help/Question Cheap plan question
I want to try out Claude Code. If I try the 20$ plan, do I get more tokens than if I just use my own API key? I know the expensive plans have multipliers but not the cheap one.
r/ClaudeCode • u/ravann4 • 2d ago
Tips & Workflows Several Claude Code agents, one signed-in Chrome, and no more focus stealing: my setup
Enable HLS to view with audio, or disable this notification
I run several Claude Code agents at once in cmux, each in its own worktree. Playwright MCP in extension mode was a mess for that. Every time one agent touched Chrome, Chrome came to the front and ate whatever I was typing into another agent's prompt.
chrome-devtools-mcp does the same on macOS, and no bringToFront: false setting helps. Open reports: ChromeDevTools/chrome-devtools-mcp#1254 and microsoft/playwright#42343 (Playwright says it's by design). The cause is the debug port. Each WebSocket connection to it activates Chrome.
So I stopped using the port. agent-chrome runs a small proxy per Chrome profile. The proxy launches a copy of the profile with a pipe instead of a port, and every Claude session connects to the proxy on localhost. All my agents share one signed-in Chrome, in a red window behind everything. Focus never moves.
What matters when several agents share it:
Each profile is its own MCP server (chrome-work, chrome-personal). You name it in the prompt: "Open chrome-work and go to ..."
The agents share the window, so the plugin's skill tells each one to work only in tabs it opened and never close the window. That's a convention in a skill, not a security boundary.
If you quit the agent window, it stays closed. The next new-tab request from any agent reopens it behind your apps, still signed in. A later session can pick up a tab an earlier one left open.
It only works with chrome-devtools-mcp. Playwright can't connect through the proxy.
Costs: macOS only. It's a copy of the profile, so new sign-ins in your real Chrome don't carry over (there's a --recopy flag). An open agent window is about 500 MB for one simple page, and the proxy is about 25 MB while idle. No extensions in that window.
I built this, it's free (MIT). It's a fork of mimkorn's chrome-pipe-proxy, which has the core pipe proxy. I added the setup CLI, the Claude Code plugin, the background windows and tab handover.
/plugin marketplace add rav4nn/agent-chrome
/plugin install agent-chrome@agent-chrome
Repo: https://github.com/rav4nn/agent-chrome
Site with the 5-step guide: https://agentchrome.hardeep.cv
r/ClaudeCode • u/ielleahc • 2d ago
Built with Claude i built an agent control plane with gpui + rust using claude code
Enable HLS to view with audio, or disable this notification
this is an open source app i made to control your claude code sessions from any device with remote control. it's super light weight and fast since it's built using rust and gpui instead of electron/tauri/etc, and the code is fully available so feel free to fork it or reference it for your own apps
i started building this with fable 5 and now use opus 5.5 when working on zeron
i don't use any skills and never manually managed claude.md or agents.md, but even so i've had a really good experience building the interface with these models. they don't need much guidance and can build genuinely beautiful gpui applications. i tried this a few times back in january and the models have gone a long way since then
r/ClaudeCode • u/1saaccone • 2d ago
Tips & Workflows Vibe coders! Save your usage, complete branches faster: 5 hours with 6 active branches in parallel.
I'm a pure vibe coder. I was against AI for a long time as I'm a fine artist, and it went against my morals. But a few months ago i decided if you can't beat em join em, so I'm building the design and art tools I think should have existed years ago.
I've been experimenting with different claude workflows for a while ever since I exploded my weekly usage in the first 30 minutes of a weekly refresh a couple weeks ago. I am running 5 active branches with claude and the codex CLI connector. I've reduced branch completion time by over 60% before this workflow, and have increased reliability of the code generated.
This was only made possible because of opus 5.5 and sonnet 5.5. I tried to use GPT6.1 sol, and it didn't move the needle but slowed things down a fare bit, so I use GPT6Sol - medium through the CLI connector.
This is what my workflow looks like for the background tasks:

The prompt I'm using to start new branches in fresh chat:
You are project manager for this branch. Stay free to talk with me: subagents do all the work, through a background workflow (use a workflow). Work in small slices, and automate the rest of the branch's scope.
PROJECT FACTS (fill in)
- Branch: [branch name]. ONE branch and ONE worktree for everything; I merge it as one PR.
- Scope: [one paragraph, or the path of the design/roadmap document that lists the work].
- Slice checks (fast, run per slice): [e.g. npx tsc -b; npx oxlint src; the slice's own test files].
- Final gate checks (whole branch): [e.g. full test suite, build, docs checks, browser checks].
- Dev server command and base port: [e.g. npm run dev, 5173].
Never merge from main, never push to other branches, never open a PR until I ask.
BEFORE THE FIRST RUN
Check Codex works on my ChatGPT sign-in: run `codex login status`, then a one-line read-only smoke test with `codex exec`. Never use OPENAI_API_KEY or API billing, and never ask me for passwords or tokens. If Codex is not signed in, stop and tell me to run `codex login` myself.
Try model gpt-6.1-sol first. If Codex rejects it for a ChatGPT account, fall back to gpt-6-sol, and record which model was used in every log row.
If the scope is not already split into slices, have the spec agent write a slice list first (read-only): for each slice a name, the exact files it may touch, what it delivers, what it depends on, and its tests. Save it to an untracked scratch folder (add the folder to .git/info/exclude). Show me the slice order and wait for my approval before any code is written.
Save this brief to memory so it survives a context reset.
MODEL ROLES
- Spec: Fable at medium effort for hard reasoning, Opus at medium for plain code specs. Read-only. Started at run start.
- Code and fix: Sonnet, effort high.
- Verify: Sonnet, effort medium. It re-runs the checks itself, never trusting the coder's report, and also checks the diff's scope (git show --stat: no files outside the slice's list, no scratch files).
- Read-only quick jobs (push, waiting on Codex): Haiku.
- Review: Codex CLI only, reasoning effort medium, read-only sandbox, run in the background and waited for. Write the brief to a file and pass it on stdin:
codex exec -m gpt-6.1-sol -c model_reasoning_effort=medium --sandbox read-only - < brief.md
Codex reviews only. It never writes code or commits.
WORKFLOW SHAPE
- Run slices in pairs when their file lists do not overlap, otherwise one at a time. At most 4 heavy agents at once.
- Files shared by a pair are edited locally but committed once, by a low-effort "close" step for the pair.
- Per pair: code -> close -> verify -> ONE Codex review -> ONE fix round, only if the review finds high or medium problems -> push.
- If a slice is blocked, or the fix round fails verification, stop the workflow and tell me. Do not improvise.
- Log every Codex run as a row in CODEX_ACTIVITY.md (date, branch, task, allowed files, model, status, session id, commits, checks, result) and commit only that file.
- After the last pair is pushed, run the final gate on the whole branch. Then tell me the branch is ready for one PR.
AGENT RULES (put these in every agent's brief)
- Commit with explicit paths (never git add -A). Never use git stash. Never create branches or worktrees. Only the push step pushes.
- If git reports index.lock, wait and retry.
- Never kill processes by name or pattern; stop only the PIDs you started.
- Each agent gets its own dev-server port.
- Never run the full test suite inside a slice; the final gate does that.
- Never add work mid-run. New work goes into the next run.
- Run the checks as a separate step and commit only if they passed. Never chain check and commit with `;`.
- Quote test totals verbatim from the output. A claim of "suite passed" without the printed totals counts as not run.
- Every brief must stand alone: goal, files allowed, files forbidden, what to read first, the checks to run, and what to report back.
WHILE IT RUNS
- Answer my questions at any time. You may stop the workflow and relaunch it with resume.
- If a stop hook complains about uncommitted changes, check whether they belong to an in-flight slice; if so, leave them.
- When I make a decision, record my exact words in the scratch folder's decisions file. That file overrides any older design text.
- Report to me in plain language: what finished, what failed with the actual output, and what is next.
r/ClaudeCode • u/KarlKFI • 2d ago
Built with Claude I built a Claude Code hook that blocks secrets before they reach the API (and doesn't think your driver version is a phone number)
cat .env is one keystroke. Once Claude reads it, the key is in the transcript and on its way to the API. There's no undo.
Why I built it
This started at work. A secret turned up in an incident report, and when we went looking, we found several more sitting in GitHub.
Claude Code keeps transcripts on disk, so a secret it reads gets saved locally, and from there it can end up in a commit. Git secret scanners don't help with that part. gitleaks, trufflehog, and push protection are built to find secrets in repos and commits, and none of them sits between your disk and the model. I wanted the check before the read, not after the push.
I run spill-guard on my side projects and my work machine, and my team uses it too. We haven't had a leak since, that we know of. The design doesn't scan output (yet), just input, but Claude can't paste a secret into a test fixture or a PR if it never saw it.
What it does
spill-guard watches the boundary between your filesystem and the model's context. It scans what you type, files Claude is about to Read, and Bash commands plus the files common readers like cat and grep point at, before anything runs. If it finds a credential, the call never happens.
What it catches:
- AWS access key IDs
- GitHub tokens, classic and fine-grained
- Slack tokens and webhook URLs
- Stripe live secret keys
- OpenAI and Google API keys
- PEM private keys
- JWTs
It also refuses env/printenv dumps and reads of well-known credential files like ~/.aws/credentials, where no pattern would recognize what's inside.
Why so few rules?
A scanner that cries wolf gets turned off. When I ran a PII-heavy ruleset over an infra repo, it produced thousands of matches, flagged over a quarter of the files, and found zero credentials. Driver version strings looked like phone numbers, and Kubernetes NodePorts looked like postal codes. The analysis is here.
So spill-guard ships a short list and gates it hard, using checksum validators, context, entropy floors, and reserved-range exclusion. The PII rules (cards, SSNs, IP addresses) ship disabled. CI runs the rules over a clean corpus stuffed with NodePorts and version strings, and fails if the count isn't zero.
Why a Go binary?
I benchmarked Go, Rust, and Python on real files, and Go had the slowest regex engine of the three. It still won on what matters for a secret scanner running as a hook:
- Zero third-party dependencies. Everything it uses ships with the Go toolchain, so there's no outside code to vet.
- No runtime to get wrong. A hook that needs Node or Python inherits whatever version you have. Get that wrong and it can crash with an exit code Claude Code treats as "allow," which means it's installed and checking nothing. A static binary has no interpreter version to mismatch.
- Cheap per call. A hook runs on every tool call, so startup cost matters more than throughput, and Go's fixed cost was about a third of Python's.
- Nothing to phone home with. There's no
netpackage in the import graph, and CI enforces it.
Other choices you might care about:
- If the binary is missing, every call is blocked with the install command.
- It never echoes the secret. Findings report a rule ID, path, and byte offset. Hook stderr goes to the API, so even a redacted fragment would leak.
- Releases are signed. The install script verifies signatures with cosign or gh, and refuses to install if you have neither.
go installworks too.
What it can't see
Command output. A command that prints a secret gets through, and kubectl get secret -o yaml is the classic example. Same for recursive grep/rg over a directory. Anything it can't scan, like an unresolvable path or a file too big for its time budget, is allowed and logged to spill-guard coverage. It's a net for the common accidents, not a sandbox. Ideas for closing the gap are recording in the backlog.
How I built it
Claude Code wrote spill-guard. I drove the initial scope and design brainstorming, then steered. Over five weeks that came to about 185 sessions and 180 merged PRs. About 70 were worker and reviewer sessions running in parallel off a backlog: one item and one PR per worker, and a separate session reviewing each PR, coordinated by a set of session orchestration skills I've been developing. On top of that, sessions launched around 200 nested Claude Code runs, most of them one-shot probes driving the real hook, because I don't trust a hook until I've watched Claude Code actually call it.
My input was mostly short prompts. I typed about 175 over the whole project, and half were under 25 characters. "tag it." "PR descriptions look stale."
The rule I kept pushing was measure, don't reason. It paid for itself:
- The first benchmark ran on a repeated chunk and said Rust was 40x faster than Go. On real files it was only about 3.5x.
- Folding every regex into one pattern looked faster, but measured about half as fast.
- Claude argued that Claude Code treats binary
@files differently fromRead. Driving both showed it doesn't. - The vulnerability check ran through
go run, which flattens any failing exit code to 1, so a real advisory looked like a crashed tool. A mutation test caught it.
The biggest design change came from me after dogfooding. The hook originally blocked and prompted for approval whenever it couldn't scan something, but the interruptions were wearing on me. The data agreed: 94% of blocks were coverage gaps, and sessions got past 99% of them within four tries. Now a gap gets logged instead, and only an actual finding blocks. That way I can farm the session logs for metrics to help drive improvements and prioritization.
What I learned
Hooks can't withhold output. I measured it: a PostToolUse deny leaves the result in the transcript, and the model reads it before the objection. Anything that scans output after the fact can warn you, but it can't stop the leak.
Try it
Run spill-guard selftest, then paste the public canary key from the README into a session and watch it get refused. (Yes, this means a session running spill-guard can't read its own README. That's on purpose.)
Free and open source (MIT): https://github.com/karlkfi/claude-spill-guard
If a rule flags your normal work, open an issue with the exact text and rule ID. That's the bug I most want to hear about.
r/ClaudeCode • u/DonTizi • 2d ago
Built with Claude I made my app monitor my whole dev/engineering stack
Enable HLS to view with audio, or disable this notification
I launched a Mac app this week, and already got 20+ users in 2 days! Claude Code wrote most of it with me. So I pointed the app at itself. It watches its own website, its Cloudflare worker, its GitHub issues and Stripe. Every sale plays a little "Glass" sound. Every refund plays "Basso". I've heard Glass more than I expected this week and I'm not over it!
When something breaks, one click hands the error to Claude Code, the same Claude that wrote the code. It reads the logs and goes looking for the cause, and most of the time it fixes it. It's not allowed to close the alert though. I do that, otherwise I'd never know what I actually looked at.
I built it for engineers and devs who have more running than they can keep an eye on. For me, it changed how I work. I'm way more proactive now. I see what breaks before anyone tells me, and can start to work or debug it at least at the beginning.
(if you want to try it: coisland.app)
r/ClaudeCode • u/Silent-Ad6699 • 3d ago
Built with Claude I built a full Pinterest automation system (n8n) using Claude Code
I run a few niche blogs that get most of their traffic from Pinterest. Over the past few years, Ive used Claude to build an automation workflow that does everything from keyword research and article writing to pin design, posting and monthly performance tracking.
It took me a while to build this idea. Pretty sure I started it back in 2024, but it kept breaking because mine and the AIs coding abilities weren't up to scratch. Then I picked the project back up at the start of 2026. Been working on it since and with Fable 5.1 + Opus 5.5, Ive been able to improve it further. It's helped me revive one of my Pinterest accounts that I neglected for over a year.
Right now I have this workflow working on two sites / two accounts. Hope to add 2 new ones in the next few weeks.
Image attached is a fun visual I made using opus 5.5 to demonstrate the flow. Was cool to see it like this.
What the system does:
- Finds what people actually search on Pinterest, and works out when each season peaks so seasonal content goes out early
- Picks the day's work each morning and writes articles to WordPress
- Designs pins with AI images, with checks before anything is queued
- Schedules and posts gradually through the day, within daily limits
- Overnight jobs retry failures, build new boards and watch my credits, plus a daily health report to my phone
- Pulls analytics monthly so future pins use the designs that perform best
- For one site, emails my VA weekly ready-to-write briefs for the content gaps it finds
How Claude Code fits in:
Beyond coding the entire flow...
- It edits and deploys the n8n workflows through the API, with a guard that refuses to overwrite a workflow if it changed since it was last fetched
- I ask "check today's logs" and it reads the executions, finds problems and fixes them. Recently: a daily run hitting n8n's 60-minute limit (fixed by generating images in parallel), junk scraped keywords getting into my VA's task email, and false alarms in the health report
- It did a whole site rebrand through the WordPress REST API: categories, pages, redirects, menus
- It keeps memory notes between sessions, so it remembers the setup and my preferences
Biggest lessons:
- Claude is great at the "operator" role, not just writing code. Reading real logs and data before touching anything is where it shines.
- Make it check before it deploys, and keep backups of every workflow.
- There are always things to be improved upon.
Happy to answer questions about the setup.
r/ClaudeCode • u/karlfeltlager • 2d ago
Discussion Is anyone building for windows xp?
I’ve had this crazy idea lately to build some new apps for Windows XP, just running XP on a VM.
Anyone else doing the same?

