An open-source project called pr-proof reports that its Claude Code skill removed 34% of noisy CodeRabbit findings in a 50-pull-request benchmark while retaining 93.5% of the benchmark’s labeled real bugs. Those are repository-reported results, not a guarantee for every team or codebase. The project validates review comments against checked-out code; it does not change CodeRabbit itself.
What pr-proof does to review comments
AI review comments are claims to check, not automatic fixes. The pr-proof repository packages three Claude Code skills that trace code execution, inspect callers and check library behavior to assess findings. One validates comments without editing anything; another can apply fixes only after the user approves them.
Validate existing comments
pr-comment-validation assigns each comment a verdict: valid, partly valid, wrong or style. It cites code evidence and makes no changes. Use it when you want a second opinion on a bot’s findings before deciding what to do.
Work through a pull request
pr-validation checks out a PR in a worktree, validates its review comments, presents the verdicts, applies fixes the user approves and replies on review threads. It is the more action-oriented path for handling comments, with approval between validation and edits.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →#1 Best Overall
Generate a separate review
pr-review writes a review of its own and has independent subagents try to disprove its findings before posting. It can also draft the review to a file instead. This is a different task from filtering CodeRabbit’s existing comments, and its benchmark results should not be conflated with the comment-filter results.
The README’s sample prompts include “are these PR comments valid?”, “handle the review comments on PR #123” and “review PR #123”. Anthropic describes skills as instructions Claude can add to its toolkit through a SKILL.md file; they can load when relevant, be invoked with /skill-name, or be shared as project skills or plugins in the Claude Code skills documentation.
Rank #2
What the CodeRabbit benchmark reports
The project reports a run on 50 real PRs drawn from Code Review Bench, a dataset whose year is not stated in the repository. The benchmark includes PRs from Sentry, Grafana, Keycloak, Discourse and Cal.com, with human-written “golden comments.” The repository says its filter received extracted comment text, file and line information, plus checked-out code; it did not see the labels. Scoring used the benchmark’s published labels, and the README says both the published results and pr-proof run used Claude Opus 4.5 as judge.
| Measure | CodeRabbit comments before filtering | After pr-proof filtering |
|---|---|---|
| Precision | 25.7% (repository-reported benchmark) | 32.9% (repository-reported benchmark) |
| Recall | 56.2% (repository-reported benchmark) | 52.6% (repository-reported benchmark) |
| F1 | 35.2% (repository-reported benchmark) | 40.4% (repository-reported benchmark) |
| Issues | 300 | 219 |
In the project’s reported run, the filter retained 72 of 77 labeled real bugs, or 93.5%, and removed 76 of 223 issues labeled as noise, or 34%. F1 increased by 5.2 percentage points, from 35.2% to 40.4%; the README reports a 95% confidence interval of +1.9 to +8.3 percentage points. These figures describe this benchmark and run, not a verified rate for CodeRabbit reviews in general.
Recommended Free Tools
Rank #3
How to interpret the results—and their limits
The result is a trade-off rather than a claim that filtering catches more bugs: recall fell from 56.2% to 52.6%, while precision and F1 rose. In practical terms, the benchmark’s filtered set contained a larger share of labeled valid findings, but it also retained a smaller share of all labeled bugs. Whether that is desirable depends on whether your priority is reducing questionable comments or seeing as many possible findings as possible.
- The labels may be incomplete. The repository notes that a finding counted as noise could be a real issue missing from the benchmark’s “golden” list. That can make measured precision understate actual quality.
- The benchmark may not predict current repositories. Its PRs are public and older than the models, so training-data leakage is possible. Results from older public code may not transfer to private, newer or differently structured projects.
- The evaluation was isolated. The headless Claude Code runs had no user settings, hooks, MCP servers, plugins, web access,
ghorcurl, and could not read the original PR discussions. A team’s configured workflow may behave differently. - Runs varied. The README reports two identical drafting runs scoring 33.5% and 28.2% F1. Confidence intervals were bootstrapped over 50 PRs, so neither those intervals nor the point estimates eliminate uncertainty about performance elsewhere.
The standalone reviewer has a different result
For pr-review, the README reports F1 of 29.8% (95% confidence interval 24.5–35.3%) versus 29.1% (25.5–33.1%) for plain Claude Code Opus 5.5. The difference was +0.7 percentage points, with a confidence interval from −3.4 to +4.8. The project characterizes this as statistically level with plain Claude Code: pr-review writes fewer comments and is more precise, but finds fewer bugs. This result concerns generating a new review, not filtering CodeRabbit comments.
Rank #4
Install pr-proof in Claude Code
The project requires Claude Code and an authenticated GitHub CLI (gh). The repository offers two installation routes:
Install the plugin
- In Claude Code, add the marketplace with
/plugin marketplace add TanayK07/pr-proof. - Install the plugin with
/plugin install pr-proof@pr-proof.
Copy the skill folders
Alternatively, copy the folders under skills/ in the repository into ~/.claude/skills/. The project is licensed under Apache-2.0.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Best Value
For a read-only check, start with pr-comment-validation and ask whether the PR comments are valid. Choose pr-validation if you want help triaging threads and proposing approved fixes. Use pr-review when the goal is to generate a review rather than evaluate an existing bot’s comments.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




