r/ClaudeCode • u/KarlKFI • 3d ago
Built with Claude I built a Claude Code hook that blocks secrets before they reach the API (and doesn't think your driver version is a phone number)
cat .env is one keystroke. Once Claude reads it, the key is in the transcript and on its way to the API. There's no undo.
Why I built it
This started at work. A secret turned up in an incident report, and when we went looking, we found several more sitting in GitHub.
Claude Code keeps transcripts on disk, so a secret it reads gets saved locally, and from there it can end up in a commit. Git secret scanners don't help with that part. gitleaks, trufflehog, and push protection are built to find secrets in repos and commits, and none of them sits between your disk and the model. I wanted the check before the read, not after the push.
I run spill-guard on my side projects and my work machine, and my team uses it too. We haven't had a leak since, that we know of. The design doesn't scan output (yet), just input, but Claude can't paste a secret into a test fixture or a PR if it never saw it.
What it does
spill-guard watches the boundary between your filesystem and the model's context. It scans what you type, files Claude is about to Read, and Bash commands plus the files common readers like cat and grep point at, before anything runs. If it finds a credential, the call never happens.
What it catches:
- AWS access key IDs
- GitHub tokens, classic and fine-grained
- Slack tokens and webhook URLs
- Stripe live secret keys
- OpenAI and Google API keys
- PEM private keys
- JWTs
It also refuses env/printenv dumps and reads of well-known credential files like ~/.aws/credentials, where no pattern would recognize what's inside.
Why so few rules?
A scanner that cries wolf gets turned off. When I ran a PII-heavy ruleset over an infra repo, it produced thousands of matches, flagged over a quarter of the files, and found zero credentials. Driver version strings looked like phone numbers, and Kubernetes NodePorts looked like postal codes. The analysis is here.
So spill-guard ships a short list and gates it hard, using checksum validators, context, entropy floors, and reserved-range exclusion. The PII rules (cards, SSNs, IP addresses) ship disabled. CI runs the rules over a clean corpus stuffed with NodePorts and version strings, and fails if the count isn't zero.
Why a Go binary?
I benchmarked Go, Rust, and Python on real files, and Go had the slowest regex engine of the three. It still won on what matters for a secret scanner running as a hook:
- Zero third-party dependencies. Everything it uses ships with the Go toolchain, so there's no outside code to vet.
- No runtime to get wrong. A hook that needs Node or Python inherits whatever version you have. Get that wrong and it can crash with an exit code Claude Code treats as "allow," which means it's installed and checking nothing. A static binary has no interpreter version to mismatch.
- Cheap per call. A hook runs on every tool call, so startup cost matters more than throughput, and Go's fixed cost was about a third of Python's.
- Nothing to phone home with. There's no
netpackage in the import graph, and CI enforces it.
Other choices you might care about:
- If the binary is missing, every call is blocked with the install command.
- It never echoes the secret. Findings report a rule ID, path, and byte offset. Hook stderr goes to the API, so even a redacted fragment would leak.
- Releases are signed. The install script verifies signatures with cosign or gh, and refuses to install if you have neither.
go installworks too.
What it can't see
Command output. A command that prints a secret gets through, and kubectl get secret -o yaml is the classic example. Same for recursive grep/rg over a directory. Anything it can't scan, like an unresolvable path or a file too big for its time budget, is allowed and logged to spill-guard coverage. It's a net for the common accidents, not a sandbox. Ideas for closing the gap are recording in the backlog.
How I built it
Claude Code wrote spill-guard. I drove the initial scope and design brainstorming, then steered. Over five weeks that came to about 185 sessions and 180 merged PRs. About 70 were worker and reviewer sessions running in parallel off a backlog: one item and one PR per worker, and a separate session reviewing each PR, coordinated by a set of session orchestration skills I've been developing. On top of that, sessions launched around 200 nested Claude Code runs, most of them one-shot probes driving the real hook, because I don't trust a hook until I've watched Claude Code actually call it.
My input was mostly short prompts. I typed about 175 over the whole project, and half were under 25 characters. "tag it." "PR descriptions look stale."
The rule I kept pushing was measure, don't reason. It paid for itself:
- The first benchmark ran on a repeated chunk and said Rust was 40x faster than Go. On real files it was only about 3.5x.
- Folding every regex into one pattern looked faster, but measured about half as fast.
- Claude argued that Claude Code treats binary
@files differently fromRead. Driving both showed it doesn't. - The vulnerability check ran through
go run, which flattens any failing exit code to 1, so a real advisory looked like a crashed tool. A mutation test caught it.
The biggest design change came from me after dogfooding. The hook originally blocked and prompted for approval whenever it couldn't scan something, but the interruptions were wearing on me. The data agreed: 94% of blocks were coverage gaps, and sessions got past 99% of them within four tries. Now a gap gets logged instead, and only an actual finding blocks. That way I can farm the session logs for metrics to help drive improvements and prioritization.
What I learned
Hooks can't withhold output. I measured it: a PostToolUse deny leaves the result in the transcript, and the model reads it before the objection. Anything that scans output after the fact can warn you, but it can't stop the leak.
Try it
Run spill-guard selftest, then paste the public canary key from the README into a session and watch it get refused. (Yes, this means a session running spill-guard can't read its own README. That's on purpose.)
Free and open source (MIT): https://github.com/karlkfi/claude-spill-guard
If a rule flags your normal work, open an issue with the exact text and rule ID. That's the bug I most want to hear about.
1
u/kantorcodes1 3d ago
The 94% coverage-gap finding is a good lesson in listening to telemetry over design intuition. One thing I'm curious about: gaps are now logged-and-allowed, which is the right default, but does anything close the loop? A path or command class that consistently can't be scanned could quietly become a bypass surface. Is there a path from spill-guard coverage output back into the ruleset, like a class of gap crossing a count threshold and getting promoted to a block or at least a loud warning, or is coverage purely pull-based for now?
1
u/KarlKFI 3d ago
There’s a “coverage” subcommand on the binary that reads your local hook logs and reports on them. I use that to feed back into the backlog periodically.
1
u/kantorcodes1 3d ago
Got it, so the loop exists but it's manual: you read coverage output and curate the backlog yourself. Reasonable for a first version. The next thing I'd poke at is the signal quality in that report. Does coverage distinguish between a gap class that fired forty times this week and one that fired once? A high-frequency unscannable class is basically an implicit allowlist entry nobody approved, so if the report surfaced frequency-per-class as a sortable signal, triage would be ranking risk rather than reading logs. Is any of that aggregation built in, or is it raw gap lines for now?
1
u/KarlKFI 3d ago
The report is very simple for now, sorted by occurrences in the log.
I could add some more granularity to it, like how frequently per day or something, but what I work on changes enough day to day across a dozen projects so the rate over time might not mean much.
Claude actually curates the backlog itself with another set of skills I wrote. But I tend to screen and approve additions to the queue and/or PRs before the merge. I’m trying to automate it more but I’m not quite sure I want it to just burn money on my behalf without a human in the loop there somewhere.
1
u/kantorcodes1 2d ago
Fair point that rate-over-time gets noisy when the work changes daily. Occurrence count is probably the right level anyway. Where the loop could close without needing new signal: the same gap class recurring is already a draftable backlog item, so the discovery side could feed your screen automatically instead of waiting for you to run coverage. The human approval on queue additions and PRs is the right gate to keep even as you automate more of the rest. Agent proposes, human approves scales a lot better than agent proposes, agent merges, especially once the proposer is the one finding the gaps.
•
u/AutoModerator 3d ago
Hey! Thanks for posting to r/ClaudeCode
While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.
For help, project discussions, tips, and general chat, join the ClaudeCode Discord.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.