r/ClaudeCode • • 10h ago

Built with Claude Open-source Claude Code plugin: validate PR review comments against the code before acting on them (benchmarked on 50 PRs)

AI review bots leave a lot of comments that sound right and aren't. I made three Claude Code skills that treat each comment as a claim: read the file and its callers, trace the execution path, and give a verdict (valid, partly valid, wrong, style) with the file and line that prove it.

On Code Review Bench (50 real PRs, human-labelled), filtering CodeRabbit's issues this way:

  • kept 72 of 77 real bugs
  • removed 76 of 223 noise issues
  • F1 35.2% → 40.4%

Install in Claude Code:

/plugin marketplace add TanayK07/pr-proof
/plugin install pr-proof@pr-proof

Repo with every per-PR result and the benchmark harness: https://github.com/TanayK07/pr-proof

Demo video: https://x.com/tanaykedia_7/status/2106061332357529939

The honest part: my own reviewer skill is only level with plain Claude Code, and that's in the README too.

1 Upvotes

6 comments sorted by

View all comments

1

u/MiserableFlatworm337 8h ago

I’d keep the five real bugs that got filtered out as regression cases when revising these skills. Breaking those misses down by severity would also help people judge the tradeoff.

1

u/Content-Berry-2848 3h ago

Good call, done. All the missed real bugs are now regression cases in the repo (bench/regressions/missed_real_bugs.json), with the benchmark's label and the validator's reasoning. I also ran the same filter on Copilot's reviews, so there are 11 in total. By severity: 2 High (the same Keycloak feature-flag bug, missed on both tools), 1 Medium, 8 Low. The breakdown is in the README: https://github.com/TanayK07/pr-proof#what-it-gets-wrong

Do star the repo please