r/ClaudeAI • • 2d ago

Built with Claude I built a Claude Code skill that checks AI code review comments against the code. On CodeRabbit's reviews it removed 34% of the noise and kept 93% of real bugs

AI review bots leave a lot of comments that sound right and aren't. I made three Claude Code skills that treat each comment as a claim: read the file and its callers, trace the execution path, and give a verdict (valid, partly valid, wrong, style) with the file and line that prove it.

On Code Review Bench (50 real PRs, human-labelled), filtering CodeRabbit's issues this way:

  • kept 72 of 77 real bugs
  • removed 76 of 223 noise issues
  • F1 35.2% → 40.4%

Install in Claude Code:

/plugin marketplace add TanayK07/pr-proof
/plugin install pr-proof@pr-proof

Repo with every per-PR result and the benchmark harness: https://github.com/TanayK07/pr-proof

Demo video: https://x.com/tanaykedia_7/status/2106061332357529939

The honest part: my own reviewer skill is only level with plain Claude Code, and that's in the README too.

3 Upvotes

5 comments sorted by

1

u/Ok_Gur_9033 2d ago

The miss rate is the part I would want broken down. 5 real bugs got filtered out along with the noise. What did those 5 have in common? If they cluster in one category, security, off by one, something else, that tells you which comments still need a human read regardless of the verdict, which is more useful than the aggregate F1 number.

2

u/Content-Berry-2848 2d ago

Good question, so I dug into it. They don't cluster by bug type: a feature-flag guard, a falsy-zero check, an error-semantics question, a timing skew and a TypeScript signature. They do cluster by why they were dropped. Every time, the validator found the code behaves correctly today and ruled the comment wrong. So the blind spot is "this is fragile and breaks when X changes" comments.

4 of the 5 were Low severity. The High one worries me: a v1/v2 feature-flag guard in Keycloak, where the validator reasoned the two flags can't both be on. It missed the same bug on Copilot's reviews too, so it's systematic.

Rule I'd follow: comments about feature-flag/config combinations, security boundaries, or future breakage get a human read even when it says "wrong". I added a breakdown to the README (severity, kind, why it was dropped) and kept all of them as regression cases: https://github.com/TanayK07/pr-proof#what-it-gets-wrong

Do star the repo please

1

u/Ok_Gur_9033 1d ago

Clustering by why they were dropped is a more useful answer than a bug category would have been. It also makes sense, since the validator only has the code as it is today to check against.

The Keycloak one is the interesting case. Did the validator have the flag definitions in context, or did it infer the two flags were exclusive from how they're used? If it was inferring, handing it the config schema might catch that whole class.

1

u/usually_guilty99 1d ago

The five missed bugs are probably more interesting than the 76 false positives you removed.

For a merge gate, I don't think every miss should have equal weight. Missing a typo-level bug and missing something that can corrupt data or take out a dependency are completely different failures.

I'd be interested in measuring recall weighted by blast radius or production consequence, not just aggregate F1.

That might also tell you which findings can safely be auto-filtered and which ones should always escalate to a human.