GitHub has launched the open ReviewBench benchmark to evaluate code review agents, in an attempt to address a fundamental problem in AI-powered review tools: the difficulty of comparing the errors they detect, those they miss, and the amount of noise they produce. The benchmark is available to researchers and teams developing code review systems, and also allows a custom system to be submitted and compared with other systems.
A benchmark built on real pull requests
GitHub designed the ReviewBench set based on an analysis of the distribution of more than 103.9 million pull requests on the platform. The set contains 219 pull requests from 187 publicly licensed open-source repositories, covering 19 programming languages, with the distribution of languages and repository sizes aligned with GitHub’s overall pattern. However, the size of the changes was reweighted in favor of medium- and large-sized, reviewable pull requests, rather than placing excessive emphasis on small changes in a single file.
What GitHub calls the “ground truth set” does not rely on a single source. Potential findings were collected from human reviewers, subsequent changes made by pull request authors, static analysis tools, and several advanced language models. After semantically overlapping findings were removed, they were evaluated according to a unified criterion stipulating that a correct finding must be correct, relevant, and non-trivial. GitHub uses the Claude Sonnet 5 model as the language evaluator, while publishing the rubric and evaluation settings to promote auditability and reproducibility.
Metrics that do not penalize the discovery of novel bugs
ReviewBench distinguishes between two families of metrics. The grounded precision, grounded recall, and grounded F1 metrics measure a system’s ability to detect issues already present in the ground-truth set, providing a direct comparison between systems. The augmented precision, augmented recall, and augmented F1 metrics also examine findings that do not match any known issue, awarding the system credit if the evaluator proves that the finding is a valid issue.
This distinction matters because a fixed set of bugs cannot necessarily be complete; a new agent may discover an issue that the benchmark’s creators did not identify. GitHub uses grounded recall as the primary indicator for comparing systems, while the augmented metrics provide additional diagnostics for each system.
An adjustable and auditable evaluation
Results can be broken down by issue severity and category, such as correctness, security, reliability, maintainability, and testing. The Fβ metric also allows the weighting between recall and precision to be changed, enabling the user to favor broader coverage or fewer, more accurate comments. Before launch, senior engineers who had not participated in building the data relabeled all findings, and the agreement rate with ReviewBench judgments reached 96.6%. GitHub says it versions the data, evaluator, and matching tool used in every evaluation.
What does it prove in practice?
GitHub used ReviewBench to evaluate successive versions of Copilot Code Review and said that the direction of improvement or decline in offline tests matched production experiments. In an experiment with a review system that combines multiple runs of different models, the benchmark predicted increases in precision, recall, and the number of comments, along with a decrease in review cost. The production A/B test moved in the same direction: the rate of comments that led to a code change increased by 8.0%, recall by 13.6%, and comment volume by 61%, while the cost per review decreased by 8.0% compared with the control. The benchmark also predicted a 227% increase in critical comments, compared with 262% in production.
The editorial reading from certi.news: the most important value of ReviewBench is not the launch of another tool, but the attempt to turn the evaluation of code review agents into a comparable and auditable process, while acknowledging that “more comments” does not necessarily mean a better review. However, the same source acknowledges that production experiments remain the final measure of user impact, and that relying in part on a language evaluator leaves an open question about the limits of consistency in automated judgment.
GitHub is making an initial research version of the benchmark available, along with the data, methodology, evaluator settings, and a self-hosted execution tool. Agents can be tested on a set of 25 pull requests, followed by a full evaluation on 219 pull requests across three rounds before requesting publication of the result on the leaderboard.