Follow the latest coverage, related explainers and connected technology stories.
GitHub has launched the open ReviewBench benchmark to evaluate code review agents using 219 pull requests from 187 repositories and 19 programming languages, with metrics that distinguish between detecting known and novel issues. The company says its offline results reflected the direction of production experiments on Copilot Code Review, while real-world user testing remains the final standard for judgment.