Free tools Windows power users keep installed
One-click scans. No signup required.
GitHub’s ReviewBench is an offline benchmark for comparing AI code review agents on a shared set of pull requests. It measures which findings reviewers catch, which they miss, and how much noise they produce—while letting teams weigh broad coverage against finding accuracy. GitHub announced it as a research preview on October 5, 2026, not as proof that one reviewer is best for every team.
Why GitHub created ReviewBench
AI code review systems can return different findings on the same change, and a reviewer that flags more potential issues may also produce more false alarms. ReviewBench is designed to make those tradeoffs comparable on a common collection of pull requests rather than relying on anecdotes or unlike-for-like demonstrations. GitHub says it also uses ReviewBench in offline evaluation of GitHub Copilot code review.
The benchmark is an evaluation signal, not a guarantee that a system will find every defect in a real repository. Its results describe performance against ReviewBench’s dataset, labels, and scoring choices.
How the dataset represents real pull requests
GitHub says it analyzed 103.9 million pull requests to characterize language, repository size, and change shape. The resulting ReviewBench corpus contains 219 public pull requests from 187 public open-source-licensed repositories, spanning 19 programming languages.
#1 Best Overall
GitHub says the language and repository-size distributions closely match its broader pull-request population. Pull-request size does not mirror that population as closely: the benchmark deliberately gives more weight to the reviewable middle and tail, retaining substantive, multi-file work instead of letting tiny, often single-file changes dominate. That makes the set more useful for testing deeper review, but it also means the benchmark should not be read as a perfectly proportionate miniature of all GitHub pull requests.
How ReviewBench builds its ground truth
A benchmark needs a reference set of findings to score what an agent catches. ReviewBench’s “golden set” combines candidate findings from four sources:
Rank #2
- Findings made by human reviewers on the pull requests.
- Issues inferred from changes authors made in follow-up commits.
- Deterministic analysis tools.
- Suggestions from multiple frontier large language models.
GitHub says these candidates are semantically deduplicated, so multiple sources identifying the same issue do not mechanically inflate the set. Each candidate is then assessed under the same rubric regardless of where it came from. A finding counts as a true positive only when it is true, relevant, and non-trivial. GitHub names Claude Sonnet 5 as the LLM grader and publishes the rubric and judge configuration with the benchmark.
This approach broadens the pool beyond issues explicitly reported in human comments, while making the rubric’s judgment consequential: a system is rewarded for surfacing meaningful, grounded findings, not merely for producing many plausible-sounding observations.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
What the scores measure
ReviewBench reports four core metrics. Precision asks how often surfaced findings are valid; recall asks what share of known findings the reviewer catches. “Grounded” metrics compare output with the golden set, while “augmented” metrics account for newly discovered issues beyond it.
| Metric | Question it answers |
|---|---|
| Grounded precision | Of the findings the agent surfaced, how many match valid findings in the golden set? |
| Grounded recall | Of the findings in the golden set, how many did the agent catch? |
| Augmented precision | How valid are surfaced findings when newly discovered issues are also taken into account? |
| Augmented recall | How much of the known issue set, including newly discovered issues, did the agent catch? |
The benchmark also breaks results down by severity—critical, medium, and low—and by categories including correctness, security, reliability, maintainability, and testing. These slices help teams examine whether an overall score hides weaknesses in an area they care about.
Rank #4
Choosing a precision-recall balance
ReviewBench includes an Fβ score that allows the balance between precision and recall to be adjusted. A recall-oriented setting suits teams that would rather review more candidate findings to avoid missing issues. A precision-oriented setting suits teams that place greater cost on noisy or invalid alerts. There is no universally best balance: the right preference depends on the team’s risk tolerance, review workflow, and the severity and category of findings it wants to prioritize.
What GitHub reports about validation
GitHub says senior engineers who had not taken part in constructing the dataset independently relabeled every ground-truth finding before release. Their true/false-positive judgments agreed with the benchmark 96.6% of the time, according to GitHub’s October 5, 2026 announcement. That is publisher-reported agreement on those judgments; it is not an independently published audit of every aspect of the benchmark.
Recommended Free Tools
Best Value
GitHub also says it checks movement in the offline benchmark against online experiments and that the offline signal has become more effective at anticipating the direction of production experiment results. These are GitHub’s validation claims, not an independent demonstration that every leaderboard ranking will transfer to every team’s codebase or review process.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to try the research preview
GitHub’s announcement describes ReviewBench as a research preview available through the ReviewBench website. Users can explore the public dataset and leaderboard or register an agent for evaluation. Registration requires a container image, configuration, and the user’s own model key.
- Explore the dataset and leaderboard on the ReviewBench website to understand the benchmark and existing submissions.
- Register an agent by providing its container image and configuration, along with your own model key.
- Run the test evaluation, which covers 25 pull requests and provides per-pull-request detail.
- Submit the final run, which evaluates all 219 pull requests in three rounds.
- Wait for maintainer review: scores remain private until a maintainer approves the submission. Publication requires a first leaderboard entry or an improvement over the current score.
Because this is a preview, availability and leaderboard contents may change.
How teams should interpret a result
Use a ReviewBench score as a structured comparison under the benchmark’s shared conditions, then look beyond the headline number. Check severity and category breakdowns, decide whether precision or recall better fits the cost of a false alarm or missed issue, and consider how representative the dataset’s deliberately substantive pull requests are of the work your team reviews. ReviewBench makes those choices more visible; it does not remove the need to judge whether the benchmark’s priorities match your own.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




