Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
42 changes: 29 additions & 13 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,9 +42,9 @@ Without shared evals for these tools, every company grades its own homework. You
| [Discourse](https://github.com/discourse/discourse) | Ruby | Forum platform |
| [Keycloak](https://github.com/keycloak/keycloak) | Java | Authentication |

Each PR has curated golden comments with severity labels (Low / Medium / High / Critical). An LLM judge matches each tool's review against the golden comments and computes precision and recall.
Each PR has curated golden comments (173 total) with severity labels (Low / Medium / High / Critical) and category tags (bug, security, concurrency, data, api, perf, test_gap, doc_defect, style, speculative). An LLM judge matches each tool's review against the golden comments using three judge models (Claude Opus 4.5, GPT-5.2, Claude Sonnet 4.5). Category-based scoring profiles (Strict / Core / All) control which issue types count toward the score, and F-beta weighting lets users prioritize recall over precision.

**Tools evaluated**: Augment, Claude Code, CodeRabbit, Codex, Cursor Bugbot, Gemini, GitHub Copilot, Graphite, Greptile, Propel, Qodo, and more. Adding a new tool takes an afternoon — fork the benchmark PRs, trigger the tool, run the pipeline.
**Tools evaluated**: Augment, Baz, Claude Code, CodeAnt, CodeRabbit, Cubic, Cursor Bugbot, Devin, Gemini, GitHub Copilot, GitLab Duo, Graphite, Greptile, KG, Kodus, Macroscope, Qodo, Sourcery, and more. Running a tool that isn't on the leaderboard takes an afternoon — fork the benchmark PRs, trigger the tool, run the pipeline. Publishing it alongside the others additionally requires meeting the [inclusion criteria](#inclusion-criteria).

> **Known limitation**: Static datasets risk training data leakage — tools may have seen these PRs during training. That's why we also run the online benchmark.

Expand All @@ -55,15 +55,16 @@ See [`offline/README.md`](offline/README.md) for setup and usage.
The online benchmark continuously samples **fresh real-world PRs from GitHub** where code review bots left comments. Because the PRs are recent, tools can't have memorized them during training.

```
GitHub Archive (BigQuery)
GitHub Search API
│
▼
┌────────┐ ┌─────────┐ ┌─────────┐ ┌────┐ ┌───────────┐
│Discover│────▶│ Enrich │────▶│ Analyze │────▶│ DB │────▶│ Dashboard │
└────────┘ └─────────┘ └─────────┘ └────┘ └───────────┘
BigQuery scan GitHub API LLM 3-step Postgres Interactive
finds bot PRs fetches full extraction & or SQLite filters &
PR context matching time series
Search API GitHub API LLM 3-step Postgres Interactive
finds merged fetches full extraction & or SQLite filters &
bot-reviewed PR context matching time series
PRs
```

**How analysis works**:
Expand All @@ -74,7 +75,7 @@ GitHub Archive (BigQuery)

**Bots tracked**: CodeRabbit, GitHub Copilot, Claude, Cursor, Augment, Codex, Gemini, Greptile, Graphite, Qodo, Propel, and others.

**Dashboard features**: Filter by language, project domain, PR type, issue severity, diff size. Track performance over time. Adjustable F-beta weighting.
**Dashboard features**: Filter by language, project domain, PR type, issue severity, diff size, engagement signals (human comments/commits after bot review), solo-bot PRs, and sample controls. Track performance over time. Adjustable F-beta weighting. See [`online/FILTERS.md`](online/FILTERS.md) for the full filter spec.

See [`online/README.md`](online/README.md) for architecture and setup.

Expand All @@ -84,14 +85,15 @@ Both benchmarks use an LLM-as-judge approach, but with different methodologies s

| | Offline | Online |
|---|---|---|
| **Ground truth** | Human-curated golden comments | Developer's post-review fixes |
| **Ground truth** | Human-curated golden comments (173, categorized) | Developer's post-review fixes |
| **Precision** | Tool comments that match a golden comment / total tool comments | Bot suggestions matched to real fixes / total suggestions |
| **Recall** | Golden comments found by the tool / total golden comments | Real fixes caught by the bot / total fixes made |
| **Recall** | Golden comments in active profile found by the tool / total golden in profile | Real fixes caught by the bot / total fixes made |
| **Scoring** | Category-based profiles (Strict/Core/All) + F-beta (0.5–3.0) | F-beta with adjustable weighting |
| **Judge input** | Golden comment + tool candidate | Full PR timeline: diff, bot comments, post-review commits |

In both cases, the judge prompt asks "do these describe the same underlying issue?" — different wording is fine, only the substance matters.

**Judge model variance**: Different LLM judges can score differently. We mitigate this by storing results per judge model and reporting which model was used. The offline benchmark has been evaluated with Claude Opus 4.5, Claude Sonnet 4.5, and GPT-5.2.
**Judge model variance**: Different LLM judges can score differently. We mitigate this by storing results per judge model and reporting which model was used. The offline benchmark has been evaluated with Claude Opus 4.5, Claude Sonnet 4.5, and GPT-5.2 — the top 5 tools are identical across all three judges, with most tools varying by at most 2 rank positions.

## Repository structure

Expand Down Expand Up @@ -135,7 +137,7 @@ uv run python -m code_review_benchmark.step2_5_dedup_candidates

# Run the LLM judge (pass dedup groups to avoid penalising duplicate candidates)
uv run python -m code_review_benchmark.step3_judge_comments \
--dedup-groups results/$(MARTIAN_MODEL)/dedup_groups.json
--dedup-groups results/${MARTIAN_MODEL}/dedup_groups.json

# View results
open analysis/benchmark_dashboard.html
Expand All @@ -161,15 +163,29 @@ uv run python main.py analyze --all
uv run python main.py dashboard
```

## Adding a new tool to the offline benchmark
## Adding a new tool

Running a tool that isn't on the leaderboard is open to anyone:

1. Fork the 50 benchmark PRs into a GitHub org where your tool is installed
2. Let the tool review each PR
3. Add the tool name to the download config and run the pipeline
4. Results appear alongside existing tools in the dashboard
4. Compare the results against the existing tools in the dashboard

See [`offline/README.md`](offline/README.md) for detailed instructions.

### Inclusion criteria

Publishing a tool on the leaderboard alongside the others has two further requirements.

**Public usage.** Roughly at least 600–1,000 reviewed public PRs, across a spread of orgs, repos, and authors. The offline benchmark is a fixed set of 50 PRs, so on its own it can't tell us whether a score reflects how a tool behaves in practice. We validate it against the online benchmark, which measures how developers respond to a tool's reviews in the wild, and that cross-check needs enough public review activity to be meaningful. Private installs aren't visible to us and can't be counted, so this is a measurement constraint rather than a judgement about a tool's overall adoption.

**Attributable reviews.** We need to be able to tell from the GitHub API that a review came from the tool rather than from a person — a bot account, a dedicated machine account, or a consistent marker in the comment body all work. Without one of those we can't separate a tool's findings from a human reviewer's comments.

For any tool we publish we also run the pipeline ourselves rather than take submitted results, so every number on the leaderboard is produced the same way.

If a tool doesn't meet these yet, the harness is public and you're welcome to run it and publish your own results.

## Contributing

We welcome contributions — new tools, better golden comments, improved judge prompts, additional datasets. Open an issue or PR.
Expand Down
29 changes: 17 additions & 12 deletions methodology/full.md
Original file line number Diff line number Diff line change
Expand Up @@ -119,11 +119,9 @@ This section describes what we've built and deployed today. The methodology sect

### Online benchmark

Each day, we collect events from code review tools using the GHArchive dump, searching for events containing the agent ID associated with each tool. We group these events into PRs.
Each day, we discover merged PRs where tracked bots left reviews, using the GitHub Search API. We restrict to merged PRs because unmerged PRs rarely have enough post-review signal to score meaningfully. For each tool, we randomly sample PRs and cap per-repo to prevent any single project from dominating the sample. High-volume bots may produce more PRs than the API can return in a single query; we handle this transparently so the sample remains representative.

We filter to projects with more than 1,000 PRs, to avoid low-volume projects where behavioral signals are likely to be noisier. We plan to revisit this threshold — smaller projects may still provide useful data, but we need to examine where they do and don't before including them.

For each tool, we randomly sample PRs each day. If the sample size required to compute a statistically significant mean would exceed 10% of the population, we compute the population mean directly by pulling all examples.
For each PR, we fetch the full context — commits, reviews, threaded discussions, and per-commit diffs — and assemble a unified chronological timeline. An LLM then performs a three-step analysis: extract what the bot suggested, extract what the developer actually changed after the review, and judge which suggestions correspond to real fixes.

We compute the following metrics for each tool:

Expand All @@ -139,19 +137,26 @@ We compute the following metrics for each tool:
- Total comments acted on

*Distributional controls:*
- Distribution of diff size, repo size, language, backend vs. frontend, and other repository and PR characteristics across tools, to identify whether tools are operating on comparable datasets
- Distribution of diff size, language, project domain, PR type, issue severity, and other repository and PR characteristics across tools, to identify whether tools are operating on comparable datasets

We plot these statistics over time to track trends in tool performance and usage patterns.

Not all PRs are equally informative. A PR where no human ever looked at the review — no comments, no follow-up commits — tells us little about whether the bot's suggestions were useful. We compute engagement signals for each PR: whether humans commented or pushed commits after the bot's first review, how many distinct reviewers participated, and how many back-and-forth rounds occurred. These signals power a set of quality filters that let us progressively restrict to higher-signal subsets. By default, we exclude bot-authored PRs (where the "developer response" is another bot) and self-reviews (the bot reviewing its own PR). Users can also restrict to solo-bot PRs — where only one bot reviewed — to simplify attribution of human follow-up.

Each of these filters is exposed individually in the dashboard, so users can combine them based on what they care about — requiring human engagement, capping per-author-per-repo contributions, setting minimum contributor diversity, and so on. Stricter filter combinations trade coverage for signal quality: fewer PRs, but more reliable scores.

### Offline benchmark

We build on the Augment/Greptile dataset. For each PR in the dataset, we:
We build on the Augment/Greptile dataset, with a curated gold set of **173 golden comments** across 50 PRs. Each golden comment has a severity level and an issue-type category.

For each PR in the dataset, we:

1. Fork a copy of the repository for each code review tool being evaluated.
2. Open a PR that includes the description from the original human-authored PR.
3. Trigger the code review tool in the forked repo on GitHub.
4. Collect all issues identified by the tool, splitting multi-issue comments into individual issues.
5. Run a judge to evaluate which tool-identified issues correspond to issues in the gold set.
5. Deduplicate candidates, grouping those that express the same underlying concern so duplicate mentions are not penalized.
6. Run a judge to evaluate which tool-identified issues correspond to issues in the gold set.

The judge uses the following prompt:

Expand All @@ -168,19 +173,19 @@ The judge uses the following prompt:
>
> Respond with ONLY a JSON object: {"reasoning": "brief explanation", "match": true/false, "confidence": 0.0-1.0}

Matches against the gold set are counted as true positives. Precision and recall are computed from these matches.
Matches against the gold set are counted as true positives. Precision and recall are computed from these matches, with category-based scoring profiles and F-beta weighting available in the dashboard. Results are evaluated with three independent judge models (Claude Opus 4.5, GPT-5.2, Claude Sonnet 4.5); the top 5 tools are identical across all three judges.

The code to reproduce this is available at: https://github.com/withmartian/code-review-benchmark

### Known limitations of the current implementation

The current implementation is deliberately minimal — a starting point we can measure and improve against. The main limitations, each addressed by a subsequent section of this document:
The current implementation is under active improvement but retains significant limitations, each addressed by a subsequent section of this document:

- Bug definitions are implicit in the gold set rather than explicit or conditioned on user preferences (§5).
- Bug definitions are semi-explicit through categories but not conditioned on user preferences (§5).
- Recall is capped by the gold set, which is itself capped by human performance (§6).
- Precision treats all non-action as false positives (§7).
- Scoring profiles reduce false positive penalties for style/speculative findings, but the gold set has sparse coverage of style issues, so some correct style findings are still penalized (§7).
- The dataset is static and based on older PRs, with contamination risk (§8).
- The judge is a single LLM with a single prompt, with no calibration against human annotations (§9).
- The judge uses three models with consistent rankings, but has not been formally calibrated against human annotations (§9).
- There is no standardized harness — we're measuring product performance, not model performance (§9).

---
Expand Down
2 changes: 1 addition & 1 deletion methodology/summary.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ The gap between measure and intent is the problem. To close it, you need two thi

Our benchmark has two pieces:

**An offline benchmark** runs every tool on the same PRs with the same bug definitions and scores them against a curated gold set. This is the measure — it lets us make controlled comparisons. The v0 builds on Augment and Greptile's published dataset and improves from there.
**An offline benchmark** runs every tool on the same PRs with the same bug definitions and scores them against a curated gold set. This is the measure — it lets us make controlled comparisons. The current version builds on Augment and Greptile's published dataset with a curated gold set of 173 golden comments across 50 PRs, each categorized (bug, security, concurrency, etc.) and scored using category-based profiles and F-beta weighting.

**An online benchmark** teaches us user intent based on how developers actually use these tools. When a developer fixes a problem flagged by a tool, they're voting that the flag was useful. These behavioral signals aren't controlled by us or by vendors — they're anchored to what actually happens.

Expand Down
2 changes: 1 addition & 1 deletion offline/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,7 @@ Open replication of the code review benchmark used by companies like [Augment](h
| [Qodo](https://www.qodo.ai/) | AI code review |
| [Sourcery](https://sourcery.ai/) | AI code review |

Adding a new tool requires forking the benchmark PRs and collecting the tool's reviews — see Steps 0 and 1 below.
Running a tool that isn't listed here requires forking the benchmark PRs and collecting the tool's reviews — see Steps 0 and 1 below. Publishing a tool on the leaderboard alongside the others additionally requires meeting the [inclusion criteria](../README.md#inclusion-criteria).

## Methodology

Expand Down
15 changes: 9 additions & 6 deletions online/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,20 +7,21 @@ The online benchmark solves this by **continuously sampling fresh PRs from GitHu
## How it works

```
GitHub Archive (BigQuery)
GitHub Search API
│
▼
┌────────┐ ┌─────────┐ ┌──────────┐ ┌─────────┐ ┌────┐ ┌───────────┐
│Discover│────▶│ Enrich │────▶│ Assemble │────▶│ Analyze │────▶│ DB │────▶│ Dashboard │
└────────┘ └─────────┘ └──────────┘ └─────────┘ └────┘ └───────────┘
BigQuery scan GitHub API Build unified LLM 3-step Postgres Interactive
finds bot PRs fetches full PR timeline extraction & or SQLite filters &
PR context matching time series
Search API GitHub API Build unified LLM 3-step Postgres Interactive
finds merged fetches full PR timeline extraction & or SQLite filters &
bot-reviewed PR context matching time series
PRs
```

### 1. Discover

A BigQuery scan of [GitHub Archive](https://www.gharchive.org/) finds PRs where tracked code review bots left comments. Sampling is deterministic (FARM_FINGERPRINT-based) so runs are reproducible. Up to 500 PRs per bot per day.
The GitHub Search API (`reviewed-by:<bot> is:merged`) finds merged PRs where tracked code review bots left reviews. For high-volume bots that exceed the 1,000-result API cap, adaptive time bisection splits the query window until each chunk fits. PRs are randomly sampled and capped per-repo to prevent any single project from dominating. Up to 500 PRs per bot per day. A BigQuery / [GitHub Archive](https://www.gharchive.org/) fallback is available via `--source bq`.

### 2. Enrich

Expand Down Expand Up @@ -60,8 +61,10 @@ The dashboard shows tool performance over time with filters for:
- **Severity**: low, medium, high, critical
- **Diff size**: min/max lines changed
- **F-beta**: adjustable weighting between precision and recall
- **Engagement**: require human engagement (comments/commits after the bot's review), solo-bot PRs, minimum reviewer count
- **Sample controls**: min scored PRs, min PRs per day, max PRs per (repo, author, bot) triple

Visualizations include time series of F-beta scores, precision/recall scatter plots, and a filterable leaderboard.
Visualizations include time series of F-beta scores, precision/recall scatter plots, and a filterable leaderboard. See [`FILTERS.md`](FILTERS.md) for the full filter spec.

## Components

Expand Down
Loading
Loading