Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
World desk6 min

ReviewBench: An Open Benchmark for AI Code Review

ReviewBench compares AI code review agents on 219 pull requests using grounded and augmented metrics. Here’s how its dataset, scoring, validation, and submission workflow work.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

ReviewBench is GitHub’s offline benchmark for comparing AI code review agents on a shared set of real pull requests. It measures how well reviewers identify known issues, how much noise they produce, and whether they surface additional valid findings. Its announced dataset contains 219 pull requests from 187 public open source licensed repositories across 19 languages. Teams can also submit their own agent through the ReviewBench service, subject to its research-preview workflow and maintainer approval.

What ReviewBench evaluates

ReviewBench gives different AI code reviewers a common workload and scoring method. Rather than treating every review comment as equally useful, it evaluates whether findings are true, relevant, and non-trivial, and distinguishes coverage from noise. GitHub introduced the benchmark on October 5, 2026, describing it as an open offline evaluation. The benchmark can help compare systems under controlled conditions; it does not by itself establish how a reviewer will affect developers in every production setting.

The dataset contains 219 pull requests drawn from 187 public repositories with open source licenses and spanning 19 languages. GitHub says it analyzed 103.9 million GitHub pull requests to characterize its workload, and describes the benchmark’s language and repository-size distributions as closely matching GitHub overall. Pull request size is deliberately weighted toward the reviewable middle and tail, however, so the benchmark contains fewer tiny single-file changes and preserves more substantive multi-file cases than a simple representative sample would.

That sampling choice matters when interpreting a score: ReviewBench emphasizes changes where a reviewer has meaningful work to do. It is not a guarantee that results transfer unchanged to a team whose pull requests are typically much smaller, larger, or concentrated in different languages or repository types.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How the gold set is built

The gold set combines candidate findings from real human reviews, issues inferred from author follow-up commits, deterministic analysis tools, and multiple frontier LLMs from different model families. GitHub says it uses multiple sources because no single reviewer can be expected to identify every worthwhile issue. Overlapping findings are semantically deduplicated and evaluated under a shared rubric. A finding counts as a true positive only when it is true, relevant, and non-trivial.

GitHub’s announcement names Claude Sonnet 5 as the LLM grader and says the rubric and judge are published. It also says the dataset, judge, and matcher are versioned for reproducibility. Those details are important for interpreting leaderboard results: a score is meaningful alongside the exact dataset, judge, matcher, and run configuration used to produce it.

Findings can also be examined by severity and category. The announcement lists critical, medium, and low severity, with categories including examples such as correctness, security, reliability, maintainability, and testing. Those examples are not presented as an exhaustive category list. Looking at severity and category can reveal differences that a raw comment count or overall score would hide.

How ReviewBench scores agents

ReviewBench reports grounded and augmented precision, recall, and F1. The distinction is whether scoring is limited to the benchmark’s known findings or also evaluates valid findings outside that set.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Metric family What it compares How to read it
Grounded precision, recall, and F1 Agent findings against the fixed, known gold-set labels. Supports a stable comparison against the same set of known findings. Grounded precision reflects noise among matched findings; grounded recall reflects coverage of known findings.
Augmented precision, recall, and F1 Known findings plus independent judgments of unmatched findings surfaced by an agent. Can give credit for a valid issue no gold-set source identified. Augmented recall’s denominator grows as new findings are discovered, so it is not a like-for-like headline comparison across systems.

GitHub says it uses grounded recall as the headline cross-system measure because it has a fixed denominator. Augmented metrics are additional diagnostics for an individual system, including whether it finds worthwhile issues beyond the original gold set. A grounded score is not proof that an agent found every possible issue; it measures performance against the fixed known set.

Precision and recall describe different trade-offs. Higher precision means fewer incorrect or low-value findings among the agent’s comments; higher recall means it catches a larger share of known issues. ReviewBench also reports an Fβ score, with beta adjustable to place more weight on recall or precision. The leaderboard can be re-ranked for different preferences, so readers should choose a weighting that reflects their tolerance for missed defects versus review noise rather than treating one ranking as universally best.

  • Severity: compare critical, medium, and low findings separately where available; a large number of low-severity comments need not mean an agent is better at catching consequential defects.
  • Category: inspect areas such as correctness, security, reliability, maintainability, or testing to see where an agent’s strengths and gaps lie.
  • Configuration and version: check that systems were evaluated against the same dataset, judge, matcher, and run configuration before treating score differences as meaningful.

What the reported validation does—and does not—show

GitHub reports 96.6% agreement between ReviewBench judgments and an independent audit by senior engineers. In that audit, the engineers judged findings as true or false positives, and their judgments were compared with ReviewBench’s. This is a publisher-reported validation result from GitHub, which owns the benchmark; it is evidence of agreement in that audit, not an independent evaluation of the benchmark as a whole or proof that the judging rubric is error-free.

GitHub also describes one internal multi-model ensemble experiment in which ReviewBench predictions aligned directionally with a later production A/B test. Relative to the production control, GitHub reports an 8.0% increase in addressed rate, a 13.6% increase in recall, a 61% increase in comment volume, and an 8.0% decrease in cost per review. For critical comments, GitHub says the benchmark predicted a 227% increase and the online experiment measured 262%.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

These are GitHub’s figures for one reported experiment, not independently replicated results or a general guarantee that higher offline scores will improve another team’s review outcomes. GitHub defines addressed rate as the percentage of Copilot code review comments that an LLM determines prompted a corresponding developer code change, using the diff, discussion thread, reactions, resolution state, and post-review code. It describes recall in that experiment as how much additional human review is still needed. GitHub says online experiments remain the ultimate measure of user impact.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How to submit an agent

GitHub’s October 5, 2026 announcement describes the ReviewBench service as a research preview. The announced workflow is to sign in with GitHub, register an agent using a container image and configuration, and provide the submitter’s model key. Participants can iterate on a 25-PR test set with per-PR details before running the full 219-PR evaluation in three rounds. ReviewBench provides the judge. Scores remain private until a maintainer reviews and approves the submission; leaderboard results are published only if they beat the agent’s current score or represent its first leaderboard entry.

  1. Open the ReviewBench website and sign in with GitHub. Confirm that the service and submission workflow are currently available, since the announcement identifies it as a research preview.
  2. Register the agent by supplying its container image, configuration, and your model key.
  3. Use the 25-PR test set and per-PR detail to refine the agent before starting a full evaluation.
  4. Run the full evaluation on all 219 pull requests in the announced three-round process.
  5. Wait for maintainer review and approval before expecting the score to appear on the leaderboard.

ReviewBench is an offline assessment, so it can help teams screen or compare reviewers without exposing a production pull request stream to every iteration. Before relying on a leaderboard placement, examine the versioned materials and the score breakdown that match the result, then test promising configurations on your own code and workflow.

Who should use ReviewBench

ReviewBench is most useful when a team wants a shared starting point for comparing AI review agents, understanding precision-versus-recall trade-offs, or inspecting what different systems catch across a fixed collection of changes. It can also help researchers reproduce an evaluation when they use the same dataset, judge, matcher, and configuration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

It is less decisive as a stand-in for a team’s own production trial. The corpus is limited to 219 pull requests and intentionally favors more reviewable changes; the common judge is itself an LLM-based evaluator; and the reported production correlation comes from one GitHub internal example. A team should treat benchmark scores as evidence for selecting candidates, then assess those candidates against its own languages, review norms, defect priorities, and acceptable comment volume.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.