Why So Many AI Code Reviewers Claim First Place: A Field Guide to Code Review Benchmarks
A leaderboard score needs a recipe. Without it, the number cannot tell you which reviewer to buy.
AxonBuild, a consultancy that runs no review tool, read three vendor posts on 16 August 2026. Each vendor declared itself first on the same benchmark, Martian's Code Review Bench, and each printed a different score for the same tool, CodeRabbit: 51.2, 30.3 and 57.5. AxonBuild's conclusion was not an accusation: "All three are probably telling the truth." The posts described different arms and snapshots of one benchmark. "Each vendor published on a day it was leading, on an arm and a snapshot that favoured it." (AxonBuild, published 28 August 2026)
This guide explains what goes into those recipes, where they change, and which measurements deserve a place in your own evaluation. One disclosure first: Mnemoverse is building a pull-request reviewer of its own, Sigma, which has no benchmark result yet and is not graded here (a note at the end says where we stand).
TL;DR
- A benchmark result is a measurement of one configuration, on one date, against one answer key. A bare "#1" carries none of that, and every #1 post we found was published by the vendor it ranks first.
- Shared repositories do not mean comparable scores. By our count of the published data, the same five repositories carry seven answer keys of 50 to 173 items, from Greptile, Augment, Martian, Factory, Tenki, Entelligence and Kodus.
- The scorer can change the leader without changing the reviewed code. Our recomputation of Martian's public data shows its 19 March deduplication correction moving first place under all three of its judges.
- Detection, false alarms, verifier reliability and developer outcomes belong in separate columns. SWR-Bench, CRJudgeBench and c-CRAB each cover part of this; none covers all four.
Why AI code review benchmarks matter now
When agents write more of the code, review becomes the constraint. In Sonar's survey of 1,149 developers, 96% said they do not fully trust that AI-generated code is functionally correct, yet only 48% always check it before committing, and 38% said reviewing AI code takes more effort than reviewing human code. Sonar calls this a "verification bottleneck". The fieldwork ran in October 2025, the answers are self-reports, and Sonar sells code-verification tools. (Sonar, 8 January 2026; survey report)
A buyer's question is practical: does this reviewer find real problems without consuming too much engineering attention? A benchmark supplies evidence toward that answer. It does not replace the question.
How to read an AI code review benchmark
An AI code review benchmark is a defined test that scores a reviewer's findings on a set of code changes against a specified reference or outcome. Three questions decide what any score means.
What does the reviewer see? Some benchmarks give the reviewer only the diff. Others give it the repository, the pull request description, CI output or earlier review rounds. In SWE-PRBench, all eight evaluated models scored lower as the context grew from the diff alone to the full repository, on a composite score that also penalises false and redundant comments; on the diff alone they detected 15–31% of the human-flagged issues. The models ran on a 100-PR sample of the 350, and it is a single-author, Python-heavy preprint, so it does not show that less context is generally better. It shows why the context policy belongs beside the score. (SWE-PRBench)
What counts as a hit? An answer key, often called golden comments, is the set of findings a benchmark treats as correct and expects a reviewer to recover. Benchmarks differ on what a match requires: the right mechanism, the right line, or both. Some give one credit per underlying issue and some one per comment; some merge duplicates and some penalise them. The metric name matters too. Greptile's original catch rate counted only whether the planted bug was found and ignored false alarms, so it cannot be compared with an F-score, which also counts precision. (Greptile) F2, a recall-weighted variant computed as 5PR/(4P+R), rewards finding more issues over posting fewer wrong ones; F1 weighs both equally.
Who judges? An LLM judge is a language model that decides whether a reviewer's finding matches an answer-key item or meets the benchmark's correctness criteria. The judge model, its prompt and its calibration against human decisions are part of the measurement. As the next sections show, swapping the judge can double a score.
The shared 50-PR set and its many answer keys
One dataset carries more of the published vendor numbers than any other we found. In July 2025 Greptile re-introduced 50 known bugs as pull requests in five open source projects: Sentry, Cal.com, Grafana, Keycloak and Discourse. It scored only whether each tool found the original bug. (Greptile) The forks have been public on GitHub since then.
Augment re-annotated the same pull requests because, in its words, "the original public dataset was invaluable, but incomplete." Our count of its published file gives 137 golden comments. Its own product led its self-run results on the new key. (dataset, Augment report)
Martian built its offline arm from that lineage. Counting its repository history, the key held 137 comments at launch in February 2026, 136 after a March correction (the published scores kept using 137 until August), and 173 after an August audit. The audit describes its change against the original 137: "remove 2 incorrect", "rewrite 25 terse comments" and "add 38 issues missing from the golden set", so that the "golden set goes from 137 to 173 comments." The pull requests did not change. (Martian PR #49)
Other publishers labelled the same repositories again: Entelligence 67 items, Tenki 122, Factory 167 in its revised key, and Kodus 95 confirmed bugs on a 30-PR subset. (Entelligence, Tenki, Factory, Kodus) These are sizes of answer keys, not counts of the bugs that exist. A tool that catches 80% of 50 known bugs has not done the same thing as a tool that catches 80% of 173. Say "answer keys", not "bugs", and never compare scores across keys.
Martian states the consequence of an incomplete key plainly: "If a tool finds a real bug that human annotators missed, it gets scored as a false positive." Its methodology also says the judge "has not been formally calibrated against human annotations," and its README warns that "static datasets risk training data leakage." (methodology, README) Those are the benchmark owner's own disclosures, and they are the clearest neutral statement of the problem.
Martian Code Review Bench: offline and online arms
Martian's offline arm has each tool review forked copies of the 50 pull requests, then three judges (Claude Opus 4.5, GPT-5.2 and Claude Sonnet 4.5) match the findings against the golden comments. It offers Strict, Core and All scoring profiles. Since the 5 August data, the dashboard's default view is Core F2; earlier files implied F1. So a historical default score changes meaning twice over, because both the key and the metric moved. (README, dashboard data)
The online arm asks a different question. It reads public pull requests where review bots commented. Its precision is the share of a bot's suggestions that developers acted on; its recall is the share of the developer's later fixes that the bot anticipated. Martian is explicit that "our denominator is 'all issues the developer fixed,' not 'all issues that exist.'" Recall here is relative to what the developer chose to fix. (Martian, issue #12) At launch Martian described the online arm as the headline metric, with "200K+ PRs" updated daily. (launch post, 26 February 2026)
Being listed is a separate matter. Martian's published inclusion rule asks for "roughly at least 600–1,000 reviewed public PRs, across a spread of orgs, repos, and authors," and Martian runs the pipeline itself "rather than take submitted results." The rule was written into the README on 25 August 2026, after several tools were already listed. (inclusion criteria) Treat it as an eligibility policy, not a quality threshold. A new or mostly private reviewer cannot appear on the online board at all.
AI code review benchmark directory
The table separates benchmark families; it does not rank products. "Third-party access" means whether someone other than the owner can run it, not whether a tool is eligible for a leaderboard. Dates and availability reflect our review on 4 October 2026.
| Benchmark | Reviewer input and reference | Scoring | Size | Third-party access | Latest verified update |
|---|---|---|---|---|---|
| Martian offline | Forked PRs; curated golden comments | Precision, recall, F-β; three judges; Strict/Core/All | 50 PRs; 173 comments | Runnable; MIT pipeline and data; listing has separate requirements | Key: August 2026; repository: September 2026 |
| Martian online | Public merged PRs; developers' later fixes extracted by an LLM | Acted-on precision; anticipated-fix recall | Owner reported 200K+ PRs at launch | MIT pipeline; new or private tools lack observable traffic | Daily; sampling changed August 2026 |
| Greptile | Re-introduced real bug; one labelled issue per PR | Catch rate; no precision | 50 PRs | Public forks; no benchmark licence | July 2025 |
| Qodo PR-Review-Bench | Merged PRs with injected issues | Precision, recall, F1 | 100 PRs; 580 issues | MIT data; judge not published | February 2026 |
| MacroscopeBench | Bug-introducing commits plus clean control commits | Harmonic mean of precision and recall; cost and latency | Vendor reports 12,000+ bugs, 1,500+ repositories | Leaderboard only; no dataset or harness release found | September 2026 |
| SWR-Bench | PR and project context; confirmed issues plus clean PRs | LLM-judged precision, recall, F1 | 500 issue-bearing and 500 clean PRs | Runnable; MIT | June 2026 |
| SWE-PRBench | Diff at fixed context levels; human review comments | Detection rate; judge validated against humans | 350 PRs | Runnable; CC BY 4.0 data, MIT harness | March 2026 |
| AACR-Bench | Cross-file repository context; expert-verified comments | Precision, recall, F1 with line matching | 200 PRs; 1,505 comments | Runnable; Apache-2.0 | August 2026 |
| c-CRAB | Review passed to a coding agent; tests derived from human reviews | Test pass rate, no LLM judge | 184 PRs; 234 tests | Runnable with per-PR Docker setup; no licence file | April 2026 |
| CRJudgeBench | A review comment with repository context; trustworthiness label | Accuracy, per-class precision and recall | 1,199 instances; 359 in the test split | Runnable; repository CC BY 4.0 | September 2026 |
| MCR-Bench | Multi-round revisions; defect-state labels | LLM-judged detection and state tracking | 2,269 tasks; five languages | Runnable; Apache-2.0 | August 2026 |
The vendor relabels of the 50 pull requests (Augment, Factory, Tenki, Entelligence, Kodus) belong to the first family, so they are not independent evidence about new code. A dataset's licence also does not replace the licences of the open source repositories it was built from.
The benchmark wars: what changed besides the tools
Three posts, three scores for one tool
The AxonBuild reading above is the clearest case. CodeRabbit's own post of 3 March printed CodeRabbit at 51.2 on the online arm's January–February window. (CodeRabbit) cubic's post, dated 25 March, printed CodeRabbit at 30.3, in seventeenth place; those figures match Martian's 19 March offline data in our recomputation. Greptile's post of 30 July, on the online arm, printed it at 57.5, fifth. (AxonBuild) None of these is a repeated measurement under one recipe.
Some vendors kept the distinctions that matter. Qodo's 15 March post announced first place on the offline arm at 64.3 F1 for "Qodo Extended", the pre-correction figure in the table below, which it described as "currently in research preview," and printed its production configuration, fourth at 47.9 F1, alongside. (Qodo) GitLab's page, which reports its own rank, says: "We don't grade our own homework." (GitLab Duo review benchmark)
A counting correction moved first place
On 19 March 2026 Martian published re-scored offline results with duplicate findings merged. Its rationale: "tools that post the same issue in both a summary comment and as inline comments would otherwise be penalised for the duplicate." (offline README) We recomputed the public dashboard file at the commits before (16 March) and after (19 March) the change. Under the Opus judge and the then-current overall F1 view:
| Entry | 16 March | 19 March |
|---|---|---|
| Qodo Extended | 64.3 (first) | 50.3 (third) |
| cubic | 41.8 (twelfth) | 61.8 (first) |
| CodeAnt | 51.7 | 21.6 |
First place changed under all three judges. The commit regenerated the judges' evaluations with the new counting rule; it did not change the reviews or the answer key, and both files were scored against the same 137 comments. This was a documented methodology correction, not evidence of misconduct, and the earlier scores were valid for their snapshot. They are not always labelled as such: CodeAnt's March post still prints 51.7 from the 16 March offline snapshot, a figure the same entry no longer holds. (CodeAnt)
A cheap judge doubled a score
For about a day, on 9–10 March, Martian's public dashboard data carried a gpt-4o-mini judge. Our recomputation finds the same Qodo output scoring 71.6 F1 under it, against 36.0 under Opus, 34.7 under Sonnet and 30.1 under GPT-5.2. The entry disappeared with the next update, without comment. (dashboard history)
The judges disagree on first place today
The methodology says "the top 5 tools are identical across all three judges." In the dashboard's default Core F2 view on 4 October, that holds for membership: the same five tools make every judge's top five. The order does not hold. Qodo Extended v2 is first under Opus (65.1, with cubic v2 at 64.9), and cubic v2 is first under Sonnet and GPT-5.2 (59.9 each). In the Core F1 view even the membership differs: Qodo Extended v2 leads under Opus (65.8) and GPT-5.2 (58.9), cubic v2 under Sonnet (58.5), and one of the five changes by judge. (dashboard data, methodology) Across the 36 headline views the dashboard offers (three judges, three profiles, four metrics), at least six different tools come first in our recomputation. This does not make the rankings random. It means the judge and the view are part of the result. If the benchmark published each judge's agreement with human annotators, readers could weigh these disagreements instead of guessing at them.
Thresholds and windows decide the online leader
Our queries to the online leaderboard on 4 October, for the 2 September–2 October window, put cubic first at 65.2 F1 under the default filters. Lowering the minimum-volume filter from 1,500 to 1,400 pull requests produced a tie with Kody at 65.1. Removing the volume thresholds put Kody ahead at 68.3, on a smaller sample. The data are sampled, so repeated queries move by tenths. Each threshold is a defensible design choice; together they decide who leads. (leaderboard API)
The windows drift as well. Between December 2025–February 2026 and July–September 2026, 10 of the 11 bots present in both windows rose by 7.7 to 23.0 F1 points in our queries. That is not an estimate of product improvement: model upgrades, tool changes and pipeline changes cannot be separated, and the online sampling source itself was corrected on 26 August. (sampling fix, PR #54) Do not compare online numbers across windows.
A page that changed under its original date
On 4 October, cubic's post printed cubic at 65.7 F1 and CodeRabbit at 60.2, under a byline of 25 March 2026. Its page metadata showed a 4 October modification, and no change note was visible in the text. AxonBuild had recorded different figures from the same address on 16 August (cubic 61.8, CodeRabbit 30.3). We could not retrieve an earlier copy of the page, and we could not establish which arm or view the current figures come from. (cubic) The lesson is to keep dated copies of the numbers you rely on, not to infer intent.
Who may edit the measuring stick?
Benchmark owners take outside contributions, and some of them concern how particular tools are scored. One merged change rewrote the online arm's prompt for extracting a bot's findings. Martian's reviewer objected that naming a specific bot "may be too targeted," the heuristic was removed, and the change merged on 7 August. No ranking effect has been shown. (PR #47) A separate proposal, still open, would let a tool's own hidden marker exclude some of its comments from precision scoring while keeping the exclusions auditable. (PR #60) Martian also says it shares "methodology and results with builders before publication." (methodology)
These are governance choices, not evidence of influence. The questions they raise are fair ones: who approves changes to the scorer, and how are the results that a change affects versioned and labelled? Martian deserves credit for answering in public: it fixed a missing reproducibility file the day an outside contributor reported it, and that contributor confirmed the published arithmetic. (issue #50)
Academic benchmarks: stronger tests, narrower conclusions
AACR-Bench shows how incomplete an original comment set can be. Its authors expanded 391 original review comments to 1,505 through an "AI-assisted, Expert-verified" process, "representing a 285% increase in issue coverage"; 80 senior engineers reviewed 2,145 candidate comments. About three quarters of the final key was proposed by models, which raises a different risk: favouring the kinds of findings models tend to produce. (AACR-Bench)
CRJudgeBench isolates the verification step: given a review comment, is it trustworthy? Its authors report that "trustworthy comments constitute 63.79% of the test set, yet every model predicts true for at least 85.52% of the instances," with Claude Opus 5 and GPT-5.5 at 94.99%. General models caught 2.31% to 20.77% of the untrustworthy comments; the authors' own fine-tuned model reached 37.69%, which is not a general-model result. The data are drawn from Python repositories, and 370 of the 435 untrustworthy examples are synthetic perturbations checked by experts. (CRJudgeBench) The study supports testing a verifier on its own. It does not give Martian's judge error rate.
Real pull requests are harder than planted bugs. One small study reports a best F1 of 0.847 on synthetic injected bugs and 0.066 on real pull requests, and much lower scores on large diffs than on small ones. It used 50 real pull requests, so treat it as indicative. (study) MCR-Bench adds multi-round review and finds models "degrading significantly as the number of interaction rounds increases." (MCR-Bench)
Measurements that are harder for a lenient judge to inflate
Four kinds of measurement depend less on a lenient judge, and each brings its own confound. No benchmark we reviewed combines them.
- Execution. c-CRAB hands each review to a coding agent and then runs tests derived from human reviews; no LLM judge decides the pass. Claude Code passed 32.1% of the 234 tests, Devin 24.8%, PR-Agent 23.1% and Codex 20.1%, and "97 out of the 234 tests were passed by at least one tool (41.5%)." The coding agent's skill now confounds the reviewer's score, and the Docker setup per pull request makes it costly. (c-CRAB)
- Clean pull requests. SWR-Bench pairs 500 pull requests that carry real issues with 500 clean ones, so false alarms become countable. In the authors' comparison of review tools, the top combination, PR-Review with Gemini-2.5-Pro, "achieved an F1 score of only 19.38%" overall. The best overall figure in the paper is 23.84% F1, from the authors' own aggregation of several reviews. Narrower slices score higher, up to 33.77% F1 on pull requests with a single expected change. The set is built from Python projects and scored by an LLM judge. (SWR-Bench) MacroscopeBench mixes clean control commits in for the same reason, in a vendor design that is also, by its own description, "the same benchmark used to build and tune Macroscope Code Review." (Macroscope)
- The verifier alone. CRJudgeBench, above, measures whether the step that checks findings can reject wrong ones.
- Real traffic. Published "accepted" or "acted-on" rates range from 7.2% to 78.13%, and they measure different things. At Mozilla and Ubisoft, 8.1% and 7.2% of generated comments were accepted by human reviewers in a 2024 study. (study) Cursor reports a 78.13% resolution rate for its own Bugbot, measured by Cursor with an LLM judge. (Cursor) Anthropic reports that "less than 1% of findings are marked incorrect" for its Code Review service, which is an internal measure, not golden-set precision. (Anthropic) Outcomes can also pull in different directions. In one industrial deployment of a review tool based on Qodo's open-source PR-Agent, 73.8% of automated comments were resolved, yet average pull request closure time rose from 5 hours 52 minutes to 8 hours 20 minutes; at Atlassian, comments from its own RovoDev reviewer led to code changes 38.70% of the time and cycle time fell. (PR-Agent deployment, RovoDev) Different definitions, deciders and populations: never rank tools across these numbers.
Use these measurements together. Keep defect detection, verifier reliability, false alarms and developer outcomes in separate columns.
The same disease as AI memory benchmarks
The tasks differ, but the measurement problems repeat the ones we documented for AI memory benchmarks.
- Incomplete or wrong answer keys. Code review keys mostly miss real issues; memory keys can contain wrong answers. Both move scores without any change to the system under test. See wrong answer keys.
- Judge dependence. The grading model and its prompt belong in the result. See judge variance and judge leniency.
- Contamination. A public code review set risks leaking into training data, a cousin of the context-window leakage that inflates memory scores. See saturation and leakage. For the shared PR set it is a risk, not a demonstrated finding. The adjacent warning is concrete: according to Martian and a later academic account, OpenAI stopped reporting SWE-bench Verified in February 2026 over flawed tests and models reproducing gold patches from memory. We could not read OpenAI's own page during this review. (Martian, academic account)
- Recipes that do not match. Different keys, contexts, metrics and windows do not produce interchangeable scores. See the comparability key.
- Showcase configurations. Keep the distinction between a research preview and the product being bought. See the settings trap.
- A floor the judge cannot inflate. For memory we lead with a measure checked against the answer key without a judge. See the measurement floor. For code review, execution and clean pull requests play that role.
A checklist for any code review benchmark claim
Before acting on a benchmark number, ask:
- Which benchmark, and which arm, offline or online?
- Which answer-key version, and when did it last change?
- Which judge model and prompt, and has the judge been calibrated against human decisions?
- Which scoring profile and metric: F1, F2, precision or catch rate?
- Which window or snapshot, and is the number still current?
- Were false alarms measured, and were clean pull requests included?
- Did a third party run it, or the vendor?
- Is the configuration the production product or a research preview?
- Can you recompute the number from public data?
- Does the post show competitors' scores from the same run, with uncertainty?
If a "#1" survives all ten, it should carry its date, arm, judge, profile, metric and window in the same sentence. If a post cannot answer these questions, treat the number as marketing. For a shorter protocol aimed at judged tables in general, use the LLM judge field manual.
Where Mnemoverse stands
Everything in this section is a vendor statement. Sigma is Mnemoverse's first applied project: a pull-request reviewer for GitHub, now in closed beta, and you can request access. Sigma has no published benchmark result, and this guide neither grades it nor places it in any table.
We are preparing our own measurement and will publish it by the rules this guide argues for: every judge and every profile, uncertainty intervals, cost and latency per review, the exact configuration, and dated addenda instead of rewritten numbers. We intend to measure separately the stage that proposes findings and the stage that checks them before they are posted, and to count false alarms on clean pull requests.
Common questions
Can you trust AI code review benchmark results?
Read a result as a measurement of one tool configuration, on one date, against one answer key, scored by one judge. A bare rank claim tells you none of that. Knowing the full recipe lets you interpret a score; it does not by itself make the score valid. Check the answer-key version, judge, counting rules, metric and window before comparing anything.
Which AI code review benchmark is best?
No single benchmark covers everything a buyer needs. Martian's offline set gives a shared reference, SWR-Bench adds clean pull requests that make false alarms countable, CRJudgeBench tests the verification step on its own, and c-CRAB scores reviews by running tests. Pick the benchmark that measures what you care about, and keep its results in their own column.
What is the difference between Martian's offline and online benchmarks?
The offline arm runs tools on 50 fixed pull requests and matches their findings against curated golden comments. The online arm reads public pull requests and measures whether developers acted on a bot's suggestions and whether the bot anticipated the fixes developers later made.
Why do AI code review tools report different scores on the same repositories?
The same five repositories, and mostly the same 50 pull requests, carry at least seven answer keys of 50 to 173 items; one of them labels a 30-PR subset. Publishers also differ in tool configuration, judge, counting rules and metric, so shared repositories do not make scores comparable.
Does an LLM judge reliably identify incorrect review comments?
Do not assume it does. In CRJudgeBench, frontier judges labelled 94.99% of test comments trustworthy when 63.79% were, and general models caught only 2.31% to 20.77% of the untrustworthy ones. The study covers Python repositories and most of its negative examples are synthetic.
How should I evaluate an AI code reviewer for my team?
Use representative pull requests from your own repositories, include clean changes, freeze the configuration, and audit both missed issues and false alarms. Report judge sensitivity, uncertainty, cost, latency and developer outcomes separately.
How we checked these numbers
Every figure marked "our recomputation" or "our queries" was produced on 4 October 2026 from public data. The offline figures come from the Martian dashboard file at the commits named in the text (feec607 and 720e1d3 for the March correction; d382be4, 11960b1 and 444626d for the one-day judge) and from the live file. The online figures come from the public leaderboard API, queried with the window and volume filters stated beside each number. Vendor pages were read on the same day; where a page could not be read, the text says so. If any figure here is corrected, the correction will be added with its date rather than replacing the original.
Sources
Benchmark records: Martian repository, methodology, offline dashboard data, online leaderboard API, Martian launch post.
Datasets and vendor reports: Greptile, Augment, Factory, Tenki, Entelligence, Kodus, Qodo, cubic, CodeAnt, GitLab, Macroscope, Cursor, Anthropic. Vendor measurements are labelled where they are used.
Research: SWR-Bench, SWE-PRBench, AACR-Bench, c-CRAB, CRJudgeBench, MCR-Bench, synthetic versus real PRs, Mozilla and Ubisoft deployment, PR-Agent deployment, RovoDev, SWE-bench Verified account.
Context and third-party readings: Sonar developer survey, AxonBuild.
Related
- AI Memory Benchmarks: A Field Guide: the same measurement problems on the memory side
- Judges, Good and Evil: how a grading prompt alone moved a memory score by about 42 points
- LLM-as-a-Judge: Bias, Leniency and the LoCoMo Number: which direction judge bias leans
- Can You Trust an LLM Judge? A Field Manual: three questions for any judged table
- How We Measure AI Memory Honestly: comparability keys and the settings trap
- BLEU vs ROUGE vs F1 vs SARI: how overlap metrics like BLEU and ROUGE differ from task metrics like F1
— Edward Izgorodin · Last updated 2026-10-04
— Mnemoverse is a persistent-memory API for AI agents. Free key: console.mnemoverse.com · Plans and limits · Docs: Getting Started
Sigma, our pull-request reviewer, is in closed beta: request access.
