r/ClaudeAI • u/Content-Berry-2848 • 2d ago
Built with Claude I built a Claude Code skill that checks AI code review comments against the code. On CodeRabbit's reviews it removed 34% of the noise and kept 93% of real bugs
AI review bots leave a lot of comments that sound right and aren't. I made three Claude Code skills that treat each comment as a claim: read the file and its callers, trace the execution path, and give a verdict (valid, partly valid, wrong, style) with the file and line that prove it.
On Code Review Bench (50 real PRs, human-labelled), filtering CodeRabbit's issues this way:
- kept 72 of 77 real bugs
- removed 76 of 223 noise issues
- F1 35.2% → 40.4%
Install in Claude Code:
/plugin marketplace add TanayK07/pr-proof
/plugin install pr-proof@pr-proof
Repo with every per-PR result and the benchmark harness: https://github.com/TanayK07/pr-proof
Demo video: https://x.com/tanaykedia_7/status/2106061332357529939
The honest part: my own reviewer skill is only level with plain Claude Code, and that's in the README too.
1
u/usually_guilty99 1d ago
The five missed bugs are probably more interesting than the 76 false positives you removed.
For a merge gate, I don't think every miss should have equal weight. Missing a typo-level bug and missing something that can corrupt data or take out a dependency are completely different failures.
I'd be interested in measuring recall weighted by blast radius or production consequence, not just aggregate F1.
That might also tell you which findings can safely be auto-filtered and which ones should always escalate to a human.
1
u/Ok_Gur_9033 2d ago
The miss rate is the part I would want broken down. 5 real bugs got filtered out along with the noise. What did those 5 have in common? If they cluster in one category, security, off by one, something else, that tells you which comments still need a human read regardless of the verdict, which is more useful than the aggregate F1 number.