GitHub opens ReviewBench to compare what AI code reviewers catch and miss
GitHub’s ReviewBench research preview compares AI code-review tools on 219 real pull requests from 187 open-source repositories. It measures valid findings and missed issues against a fixed reference set, while separately crediting valid new discoveries. Published judging rules, severity filters and a runner let participants test their own agents; leaderboard results require maintainer review before becoming public.
Artificial Intelligence··Morning
Real code changes anchor the comparison
GitHub, the software development platform, opened ReviewBench on October 5 as a research preview for comparing AI code-review agents. Its shared dataset contains 219 real pull requests, proposed changes to repository code, drawn from 187 open-source repositories in 19 programming languages. Participants can test how their reviewers identify issues on the same cases, using the published dataset, judging method and runner. GitHub says it analyzed 103.9 million pull requests to establish distributions for language, repository size and the shape of changes. The selected cases give less weight to tiny single-file edits and include more substantive changes spanning multiple files.[1], [2]
Known issues and new discoveries receive separate scores
The reference findings combine human reviews, authors’ subsequent fixes, deterministic analysis tools and several model families. Findings describing the same issue are merged. To qualify, an issue must be true, relevant and nontrivial. Claude Sonnet 5 serves as the judge under a published rubric. GitHub also describes an internal audit in which senior engineers who had not built the dataset independently relabeled reference findings. Categories include correctness, security, reliability, maintainability and testing, and findings also receive severity labels.[1], [2]
Precision measures how many of a reviewer’s reported issues are valid; recall measures how many known valid issues it finds. F1 gives both equal weight. Grounded metrics use a fixed reference set, while augmented metrics credit valid discoveries outside it. GitHub keeps fixed-reference recall as its main cross-agent comparison because the augmented recall denominator varies with each agent’s new findings. Users can change the F-beta weighting to favor broader coverage or fewer false alarms and filter results by severity or issue category.[1], [2]
Submitted results stay private until maintainer review
Participants register their reviewer using a container image, configuration and their own model key. A 25-pull-request test set precedes three rounds across all 219 requests. Submitted results remain private until a maintainer reviews them. An agent’s first entry, or a result exceeding its existing public score, can then appear on the leaderboard. GitHub says it has compared offline evaluation signals with Copilot production experiments; this is the developer’s account of validation.[1], [2]