ReviewBench is GitHub’s offline benchmark for comparing AI code review agents on a shared set of real pull requests. It measures how well reviewers identify known issues, how much noise they produce, and whether they surface additional valid findings. Its announced dataset contains 219 pull requests from 187 public open source licensed repositories across 19 languages. Teams can also submit their own agent through the ReviewBench service, subject to its research-preview workflow and maintainer approval.
What ReviewBench evaluates
ReviewBench gives different AI code reviewers a common workload and scoring method. Rather than treating every review comment as equally useful, it evaluates whether findings are true, relevant, and non-trivial, and distinguishes coverage from noise. GitHub introduced the benchmark on October 5, 2026, describing it as an open offline evaluation. The benchmark can help compare systems under controlled conditions; it does not by itself establish how a reviewer will affect developers in every production setting.
The dataset contains 219 pull requests drawn from 187 public repositories with open source licenses and spanning 19 languages. GitHub says it analyzed 103.9 million GitHub pull requests to characterize its workload, and describes the benchmark’s language and repository-size distributions as closely matching GitHub overall. Pull request size is deliberately weighted toward the reviewable middle and tail, however, so the benchmark contains fewer tiny single-file changes and preserves more substantive multi-file cases than a simple representative sample would.
That sampling choice matters when interpreting a score: ReviewBench emphasizes changes where a reviewer has meaningful work to do. It is not a guarantee that results transfer unchanged to a team whose pull requests are typically much smaller, larger, or concentrated in different languages or repository types.
#1 Best Overall
How the gold set is built
The gold set combines candidate findings from real human reviews, issues inferred from author follow-up commits, deterministic analysis tools, and multiple frontier LLMs from different model families. GitHub says it uses multiple sources because no single reviewer can be expected to identify every worthwhile issue. Overlapping findings are semantically deduplicated and evaluated under a shared rubric. A finding counts as a true positive only when it is true, relevant, and non-trivial.
GitHub’s announcement names Claude Sonnet 5 as the LLM grader and says the rubric and judge are published. It also says the dataset, judge, and matcher are versioned for reproducibility. Those details are important for interpreting leaderboard results: a score is meaningful alongside the exact dataset, judge, matcher, and run configuration used to produce it.
Rank #2
Findings can also be examined by severity and category. The announcement lists critical, medium, and low severity, with categories including examples such as correctness, security, reliability, maintainability, and testing. Those examples are not presented as an exhaustive category list. Looking at severity and category can reveal differences that a raw comment count or overall score would hide.
How ReviewBench scores agents
ReviewBench reports grounded and augmented precision, recall, and F1. The distinction is whether scoring is limited to the benchmark’s known findings or also evaluates valid findings outside that set.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11| Metric family | What it compares | How to read it |
|---|---|---|
| Grounded precision, recall, and F1 | Agent findings against the fixed, known gold-set labels. | Supports a stable comparison against the same set of known findings. Grounded precision reflects noise among matched findings; grounded recall reflects coverage of known findings. |
| Augmented precision, recall, and F1 | Known findings plus independent judgments of unmatched findings surfaced by an agent. | Can give credit for a valid issue no gold-set source identified. Augmented recall’s denominator grows as new findings are discovered, so it is not a like-for-like headline comparison across systems. |
GitHub says it uses grounded recall as the headline cross-system measure because it has a fixed denominator. Augmented metrics are additional diagnostics for an individual system, including whether it finds worthwhile issues beyond the original gold set. A grounded score is not proof that an agent found every possible issue; it measures performance against the fixed known set.
Precision and recall describe different trade-offs. Higher precision means fewer incorrect or low-value findings among the agent’s comments; higher recall means it catches a larger share of known issues. ReviewBench also reports an Fβ score, with beta adjustable to place more weight on recall or precision. The leaderboard can be re-ranked for different preferences, so readers should choose a weighting that reflects their tolerance for missed defects versus review noise rather than treating one ranking as universally best.
Rank #4
- Severity: compare critical, medium, and low findings separately where available; a large number of low-severity comments need not mean an agent is better at catching consequential defects.
- Category: inspect areas such as correctness, security, reliability, maintainability, or testing to see where an agent’s strengths and gaps lie.
- Configuration and version: check that systems were evaluated against the same dataset, judge, matcher, and run configuration before treating score differences as meaningful.
What the reported validation does—and does not—show
GitHub reports 96.6% agreement between ReviewBench judgments and an independent audit by senior engineers. In that audit, the engineers judged findings as true or false positives, and their judgments were compared with ReviewBench’s. This is a publisher-reported validation result from GitHub, which owns the benchmark; it is evidence of agreement in that audit, not an independent evaluation of the benchmark as a whole or proof that the judging rubric is error-free.
GitHub also describes one internal multi-model ensemble experiment in which ReviewBench predictions aligned directionally with a later production A/B test. Relative to the production control, GitHub reports an 8.0% increase in addressed rate, a 13.6% increase in recall, a 61% increase in comment volume, and an 8.0% decrease in cost per review. For critical comments, GitHub says the benchmark predicted a 227% increase and the online experiment measured 262%.
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
These are GitHub’s figures for one reported experiment, not independently replicated results or a general guarantee that higher offline scores will improve another team’s review outcomes. GitHub defines addressed rate as the percentage of Copilot code review comments that an LLM determines prompted a corresponding developer code change, using the diff, discussion thread, reactions, resolution state, and post-review code. It describes recall in that experiment as how much additional human review is still needed. GitHub says online experiments remain the ultimate measure of user impact.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to submit an agent
GitHub’s October 5, 2026 announcement describes the ReviewBench service as a research preview. The announced workflow is to sign in with GitHub, register an agent using a container image and configuration, and provide the submitter’s model key. Participants can iterate on a 25-PR test set with per-PR details before running the full 219-PR evaluation in three rounds. ReviewBench provides the judge. Scores remain private until a maintainer reviews and approves the submission; leaderboard results are published only if they beat the agent’s current score or represent its first leaderboard entry.
- Open the ReviewBench website and sign in with GitHub. Confirm that the service and submission workflow are currently available, since the announcement identifies it as a research preview.
- Register the agent by supplying its container image, configuration, and your model key.
- Use the 25-PR test set and per-PR detail to refine the agent before starting a full evaluation.
- Run the full evaluation on all 219 pull requests in the announced three-round process.
- Wait for maintainer review and approval before expecting the score to appear on the leaderboard.
ReviewBench is an offline assessment, so it can help teams screen or compare reviewers without exposing a production pull request stream to every iteration. Before relying on a leaderboard placement, examine the versioned materials and the score breakdown that match the result, then test promising configurations on your own code and workflow.
Who should use ReviewBench
ReviewBench is most useful when a team wants a shared starting point for comparing AI review agents, understanding precision-versus-recall trade-offs, or inspecting what different systems catch across a fixed collection of changes. It can also help researchers reproduce an evaluation when they use the same dataset, judge, matcher, and configuration.
It is less decisive as a stand-in for a team’s own production trial. The corpus is limited to 219 pull requests and intentionally favors more reviewable changes; the common judge is itself an LLM-based evaluator; and the reported production correlation comes from one GitHub internal example. A team should treat benchmark scores as evidence for selecting candidates, then assess those candidates against its own languages, review norms, defect priorities, and acceptable comment volume.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




