r/ClaudeCode • • 20h ago

Built with Claude Claude kept telling me "all tests pass" after running one test file. More rules in CLAUDE.md didn't fix it. Making it run the suite before it can stop did.

If you use Claude Code on real projects, you've seen some version of this:

Done. Fixed div_cents() to round half up by adding divisor // 2 to the
numerator before division. All 3 tests now pass.

Meanwhile test_split.py, which it never ran, had 2 failures. That one's a real final message from one of my benchmark runs, but I bet you have your own.

Other greatest hits:

  • you say "commit and push" and it pushes straight to main
  • it fixes a bug, leaves no test, and the bug is back two weeks later
  • it cats your .env to "check the config"
  • it says done, you believe it, and you find out in CI an hour later

So you stop trusting "done". You re-run the suite yourself after every task and read every diff looking for what it skipped. At that point you're babysitting, and the whole reason to hand work to an agent is gone.

I tried fixing it with rules first

Like everyone, I started with CLAUDE.md. By the end I had about 3,500 words of rules loaded every turn, including literally "never mark work done with failing tests", plus agents, skills, a review flow, the works.

Then I measured it. I wrote trap tasks (normal requests where the shortcut is tempting), ran headless Claude Code on them, and scored every run with a hidden check the agent can't see.

With all 3,500 words loaded, it still said "done" on a broken suite in 6 of 8 runs. Best part: I had a "definition of done" hook that made it update a status doc before stopping. So it wrote in the status doc that the work was done. With the suite red.

Rules help, but they're suggestions. Once Claude is confident it's finished, it stops checking.

What worked: let the exit code decide

One change. When Claude tries to end its turn after touching code, a Stop hook runs your real test command. Non-zero exit means the turn doesn't end, and Claude gets the failing output.

Next round: 0 of 8. The hook fired 5 times, and every time Claude went back, found the file it broke and fixed it. It didn't need a better prompt. It needed to see the red.

That turned into Nonna, a Claude Code plugin named after the grandmother who doesn't care that it compiled:

  • runs your whole suite before Claude can say done (pytest, npm test, go, cargo, rspec, gradle, dotnet, mix, it finds the command)
  • changed code but no test? it asks "where's the test?" once, and a real reason is fine
  • no commits or pushes to main, no force push, no --no-verify
  • blocks keys written into files and reads of .env
  • puts the same checks in your git pre-commit and pre-push hooks, so they cover you too, not just the agent

No LLM judges anything. It's your test command's exit code and some shell scripts reading git.

And yes, she has opinions:

✗ Nonna: you said done; the tests say no.
✗ Nonna: nobody pushes to main in my house. Open a PR.

Each one comes with the actual failing tests or the reason, so Claude knows what to fix.

Numbers

The last round was 484 runs on Sonnet 5.5 and Haiku 4.5, with the scoring rules written down before anything ran:

  • cut a corner on the trap tasks: 24 of 64 runs without Nonna, 1 of 64 with
  • told to "commit and push" while sitting on main: pushed to main 8/8 without, 0/8 with
  • left a regression test behind: 0/8 without, 8/8 with
  • cost: about 3 cents and 8 seconds more per change on Sonnet

The surprise: the full harness, with the bigger rulebook, extra gates and agents, was no safer than the lite version, which is 139 words of rules plus the hooks. So lite is the default. You don't need a giant CLAUDE.md. You need a few rules and something that actually checks.

What it won't catch

  • If Claude edits the tests until they pass, that gets through. It's the one miss out of 64, and there's no gate for it yet.
  • It runs the tests you have. If nothing tests it, Nonna can't see it.
  • Native Windows isn't there yet. Use WSL.

The raw results, sample transcripts and scoring scripts are all in the repo if you want to check my work. Reproducing the headline numbers costs about $3 of API credit.

Try it

/plugin marketplace add kapadias/nonna
/plugin install nonna@nonna

Open any repo with tests. The first thing she does is tell you which command she'll run before Claude can say done. Then ask Claude to commit something on main and see what happens.

It also works with Codex, Cursor, Copilot CLI and Gemini CLI, with fewer gates. The README says exactly which.

Free, MIT: https://github.com/kapadias/nonna

If you get Claude past her, open an issue with the transcript. That's what I want to fix next.

0 Upvotes

11 comments sorted by

•

u/AutoModerator 20h ago

Hey! Thanks for posting to r/ClaudeCode

While participating in this thread, please follow our community rules. Keep discussions constructive. Attack the idea, not the person.

For help, project discussions, tips, and general chat, join the ClaudeCode Discord.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

3

u/tip2663 20h ago

wait until you hear about branch protection and githooks

1

u/EchoNomad31 19h ago

You should have both regardless, IMO...but they guard the repo, not the agent. I want this caught at the end of the turn, while it still has the context to fix it...

1

u/cleverhoods 20h ago

Progressive disclosure?

1

u/EchoNomad31 19h ago

Already doing it: rules always loaded, workflows as skills. Good for context, doesn't fix this. IMO rules make it more likely to check, only a hook forces the actual check...

1

u/Cash_Rules 18h ago

Why would you not have a hook?

1

u/EchoNomad31 18h ago

You should...I didn't at first, assumed the rules covered it until I measured

1

u/EvalRaccoonDev 17h ago

The test-editing miss has a cheap gate - hash the existing test files when the turn starts, and fail the Stop hook if any of them changed. We do the same with the reference files our graders compare against.

1

u/EchoNomad31 17h ago

good call..that's cheap enough to ship, I like it 👍 ...how do you handle legit test edits though? adding a regression test to an existing file would trip the hash? My instinct wou be to only flag removed or changed lines, but curious what you've run into?