Skip to content

Benchmark report

Entelligence Code Review Benchmark 2026

Entelligence was benchmarked against 7 leading AI code review tools (CodeRabbit, Greptile, Copilot, Graphite, Codex, Claude, and Bugbot) on real production bugs.

Reviewers
8
Production bugs
67
Repositories
5
The 2026 AI Code Review Benchmark

Overview

Introduction

AI code review tools promise to catch bugs before production. But do they actually work?

We tested 8 AI code reviewers in total: Entelligence, CodeRabbit, Greptile, Copilot, Graphite, Claude, Codex, and Bugbot on 67 real production bugs from open-source projects. Each bug shipped to production broke something and required a fix in a later commit.

The bugs included race conditions, security vulnerabilities, breaking API changes, and logic errors across five different repositories: Cal.com (TypeScript), Sentry (Python), Discourse (Ruby), Keycloak (Java), and Grafana (Go).

Results

Results (F1 scores)

50403020100
47.2%
Entelligence
Codex logo
45.4%
Codex
Claude logo
42.8%
Claude
Bugbot logo
39.4%
Bugbot
Greptile logo
36.9%
Greptile
CodeRabbit logo
33.0%
CodeRabbit
Copilot logo
22.6%
Copilot
Graphite logo
13.4%
Graphite

Full results table

RankReviewerF1 ScoreFound/Golden
1Entelligence47.2%30/67
2Codex45.4%27/67
3Claude42.8%29/67
4Bugbot39.4%31/67
5Greptile36.9%31/67
6CodeRabbit33.0%32/65
7Copilot22.6%23/67
8Graphite13.4%5/67

Evidence

See The Proof

Entelligence1 comments
1
**Correctness**: The `createEvent` signature change (line 101) breaks the `Calendar` interface contract by introducing a required `credentialId` parameter. Make `credentialId` optional to maintain compatibility with the interface.
CodeRabbit1 comments
1
View comment on GitHub⭐ Golden Comment
_⚠️ Potential issue_ | _🟠 Major_ **Logic bug: calendar lookup is ineffective.** When `externalCalendarId` is truthy, it's returned directly (correct). But when it's falsy, the `.find()` compares `cal.externalId === externalCalendarId` where `externalCalendarId` is falsy (e.g., `null` or `undefined`), which won't match any calendar with a valid `externalId`. This makes the fallback branch ineffective. If the intent is to find a calendar by some other criteria when `externalCalendarId` is not provided, the logic needs to be revised. If the intent is to just use `externalCalendarId` when provided and otherwise leave it undefined, then the `.find()` clause is dead code.

How we tested

Methodology

Building a Golden Dataset for Code Review Evaluation

Benchmarking AI code reviewers requires a ground truth: a set of golden comments representing what a thorough, context-aware review should look like. Catch rate alone - whether a tool flags a known bug—doesn't capture comment quality, relevance, or actionability. We set out to build a golden dataset that could measure these dimensions.

An Ensemble Approach to Ground Truth

Rather than relying on a single model or a single human's judgment, we built our golden dataset using an ensemble of frontier LLMs: Claude Opus, Gemini 2.5, and GPT-5. Each model received identical context for every pull request diff - including the inbound and outbound dependencies of the changed code. We extracted these dependencies using Language Server Protocol (LSP), which allowed us to programmatically trace function calls, imports, and type relationships across the codebase. This dependency context is critical; it's what separates superficial line-by-line feedback from reviews that understand how a change ripples through a codebase.

We prompted each model independently to generate review comments. This gave us three distinct perspectives on every diff, each shaped by the model's own reasoning patterns and training.

Consensus Through Voting

Three models means three (often different) sets of comments. To distill these into a single golden dataset, we implemented a majority voting process. A separate LLM call - GPT-4o - acted as the arbiter, evaluating which comments across the three sources addressed the same issues and selecting those with majority agreement. Comments that only one model flagged were discarded; comments that two or three models independently identified became candidates for the golden set.

This voting mechanism serves two purposes: it filters out model-specific hallucinations or stylistic quirks, and it surfaces issues that multiple sophisticated models agree are worth flagging. If Opus, Gemini, and GPT-5 all independently identify the same problem, that's a strong signal it belongs in the ground truth.

Human Validation

LLM consensus isn't infallible. To ensure quality, reviewers at our company manually examined the voting-selected comments. This human validation step caught edge cases where models agreed on something incorrect or where the voting process produced ambiguous results. The final golden dataset represents the intersection of multi-model consensus and human expert judgment.

Through this process, we generated 67 golden comments. We used the same repositories featured in Greptile's benchmark - Sentry, Cal.com, Grafana, Keycloak, and Discourse - so our results are directly comparable to existing industry benchmarks. The full dataset, including all golden comments, is available in our repository.

Evaluation Methodology

With our golden dataset established, we ran the benchmark. We cloned each repository and opened pull requests mirroring the original bug-introducing commits. We triggered each AI code review tool (Codex, Greptile, CodeRabbit, Entelligence, Graphite, Copilot, Claude Code, and Bugbot) on these PRs under default settings, excluding dependency bot PRs to focus on meaningful code changes.

For each tool, we collected every review comment and compared it against our golden comments using Claude Sonnet 4.5 as a semantic similarity judge. The LLM evaluated whether each AI-generated comment captured the same issue and intent as the corresponding golden comment.

Analysis

Combined Evaluation Analysis

1. Executive Summary

Aggregate Metrics (All Repositories)

RankReviewerF1 ScoreRecallPrecisionFound/Golden
1Entelligence47.2%44.8%50.0%30/67
2Codex45.4%40.3%51.9%27/67
3Claude42.8%43.3%42.3%29/67
4Bugbot39.4%46.3%34.4%31/67
5Greptile36.9%46.3%30.7%31/67
6CodeRabbit33.0%49.2%24.8%32/65
7Copilot22.6%34.3%16.8%23/67
8Graphite13.4%7.5%66.7%5/67

* Note: CodeRabbit was evaluated on 65 total PRs (14 out of 16 in Cal.com) due to PR/dataset availability during testing. All other tools were evaluated on the full 67 PRs.

2. Per-Repository Performance

entelligence Performance by Repository

RankRepositoryF1RecallPrecisionFound/Golden
1Cal.com68.1%81.2%58.6%13/16
2Grafana50.0%33.3%100.0%3/9
3Sentry45.5%29.4%100.0%5/17
4Keycloak38.5%33.3%45.5%4/12
5Discourse30.3%38.5%25.0%5/13
Run the 2026 Benchmark on Your Own Repository Right Now — Connect & Run NowRun the 2026 Benchmark on Your Own Repository Right Now — Connect & Run Now

TypeScript

Cal.com - Issue Detection Matrix

Summary Metrics

RankReviewerF1 ScoreRecallPrecisionFound/Golden
1Entelligence68.1%81.2%58.6%13/16
2Codex58.8%56.2%61.5%9/16
3Claude52.2%50.0%54.5%8/16
4Bugbot50.0%68.8%39.3%11/16
5Greptile47.4%56.2%40.9%9/16
6CodeRabbit46.0%78.6%32.6%11/14
7Copilot37.0%62.5%26.3%10/16
8Graphite0.0%0.0%0.0%0/16

* Note: CodeRabbit was evaluated on 65 total PRs (14 out of 16 in Cal.com) due to PR/dataset availability during testing. All other tools were evaluated on the full 67 PRs.

Per-PR Issue Detection Matrix

* The ↗ symbol indicates the tool successfully scanned and commented on the PR. It does not indicate whether the tool successfully found the specific golden bug. Please refer to the summary tables for exact detection rates.

Cal.com Issues (16 total)

Tip: Click any ✓/✗ to view the PR evidence.

PRBug DescriptionSeverityEntelligenceClaudeCodexCodeRabbitGreptileCopilotGraphiteBugbot
1Critical
2High
3High
4High
5High
6Critical
7Critical
8Critical
9High
10High
11High
12Critical
13High
14High
15Critical
16High

Key Findings for Cal.com

entelligence Strengths
  • Strong critical issue detection (5/6 = 83%)
  • Catches complex OAuth token handling issues
  • Identifies async/await problems (forEach with async)
  • Good at null/undefined access patterns
entelligence Weaknesses
  • Generates more false positives than codex (12 vs 5)
  • Missed timezone-related issue (organizerTimeZone)
  • Some false positives are valid issues not in golden set

Ruby

Discourse – Issue Detection Matrix

13 Golden Issues

Summary Metrics

RankReviewerF1 ScoreRecallPrecisionFound/Golden
1Bugbot44.4%46.2%42.9%6/13
2Codex38.1%30.8%50.0%4/13
3Greptile35.7%46.2%29.2%6/13
4Entelligence30.3%38.5%25.0%5/13
5Claude26.4%30.8%23.1%4/13
6CodeRabbit25.9%53.8%17.0%7/13
7Graphite25.0%15.4%66.7%2/13
8Copilot11.0%15.4%8.6%2/13

Per-PR Issue Detection Matrix

* The ↗ symbol indicates the tool successfully scanned and commented on the PR. It does not indicate whether the tool successfully found the specific golden bug. Please refer to the summary tables for exact detection rates.

Discourse Issues (13 total)

Tip: Click any ✓/✗ to view the PR evidence.

PRBug DescriptionSeverityEntelligenceClaudeCodexCodeRabbitGreptileCopilotGraphiteBugbot
1Critical
2High
3Critical
4High
5Critical
6Critical
7Critical
8Critical
9Critical
10Critical
11High
12High
13Critical

Key Findings for Discourse

entelligence Strengths
  • Catches SQL injection vulnerability (security-critical)
  • Identifies null pointer exceptions
  • Good at localization/i18n issues
entelligence Weaknesses
  • Missed Ruby syntax errors
  • Lower recall than CodeRabbit and Greptile on this codebase (Ruby)
  • Misses some nil dereference patterns
  • CSS/theme-related issues completely missed

Go

Grafana Repository (9 Golden Issues)

Grafana – Issue Detection Matrix

Summary Metrics

RankReviewerF1 ScoreRecallPrecisionFound/Golden
1Codex64.5%66.7%62.5%6/9
2Claude56.3%55.6%57.1%5/9
3Entelligence50.0%33.3%100.0%3/9
4CodeRabbit44.4%33.3%66.7%3/9
5Graphite44.4%33.3%66.7%3/9
6Bugbot33.8%44.4%27.3%4/9
7Greptile31.6%66.7%20.7%6/9
8Copilot23.5%33.3%18.2%3/9

Per-PR Issue Detection Matrix

* The ↗ symbol indicates the tool successfully scanned and commented on the PR. It does not indicate whether the tool successfully found the specific golden bug. Please refer to the summary tables for exact detection rates.

Grafana Issues (9 total)

Tip: Click any ✓/✗ to view the PR evidence.

PRBug DescriptionSeverityEntelligenceClaudeCodexCodeRabbitGreptileCopilotGraphiteBugbot
1High
2High
3High
4Critical
5High
6High
7High
8Critical
9High

Key Findings for Grafana

entelligence Strengths
  • Perfect precision (100%) - no false positives
  • Catches initialization nil issues
  • Good at identifying parameter mismatches
entelligence Weaknesses
  • Low recall (33.3%) on Go codebase
  • Misses race condition issues (0/3)
  • Misses stub function issues

Java

Keycloak Repository (12 Golden Issues)

Keycloak – Issue Detection Matrix

Summary Metrics

RankReviewerF1 ScoreRecallPrecisionFound/Golden
1Greptile52.6%41.7%71.4%5/12
2Claude45.5%41.7%50.0%5/12
3Entelligence38.5%33.3%45.5%4/12
4Bugbot24.0%25.0%23.1%3/12
5CodeRabbit22.2%41.7%15.2%5/12
6Copilot18.7%25.0%15.0%3/12
7Codex18.2%16.7%20.0%2/12
8Graphite0.0%0.0%0.0%0/12

Per-PR Issue Detection Matrix

* The ↗ symbol indicates the tool successfully scanned and commented on the PR. It does not indicate whether the tool successfully found the specific golden bug. Please refer to the summary tables for exact detection rates.

Keycloak Issues (12 total)

Tip: Click any ✓/✗ to view the PR evidence.

PRBug DescriptionSeverityEntelligenceClaudeCodexCodeRabbitGreptileCopilotGraphiteBugbot
1Critical
2Critical
3High
4High
5Critical
6High
7High
8High
9Critical
10High
11High
12Critical

Key Findings for Keycloak

entelligence Strengths
  • Catches HTML sanitizer issues (validation order, case sensitivity)
  • Identifies authorization ID mismatches
  • Reasonable precision (36.4%)
entelligence Weaknesses
  • Low critical issue detection (20%)
  • Missed passkey authentication bugs
  • Missed duplicate null check bug

Python

Sentry - Issue Detection Matrix

Summary Metrics

RankReviewerF1 ScoreRecallPrecisionFound/Golden
1Codex46.2%35.3%66.7%6/17
2Entelligence45.5%29.4%100.0%5/17
3CodeRabbit37.5%35.3%40.0%6/17
4Bugbot35.0%41.2%30.4%7/17
5Claude34.3%41.2%29.4%7/17
6Greptile24.5%29.4%21.1%5/17
7Copilot19.7%29.4%14.8%5/17
8Graphite0.0%0.0%0.0%0/17

Per-PR Issue Detection Matrix

* The ↗ symbol indicates the tool successfully scanned and commented on the PR. It does not indicate whether the tool successfully found the specific golden bug. Please refer to the summary tables for exact detection rates.

Sentry Issues (17 total)

Tip: Click any ✓/✗ to view the PR evidence.

PRBug DescriptionSeverityEntelligenceClaudeCodexCodeRabbitGreptileCopilotGraphiteBugbot
1Critical
2Critical
3Critical
4Critical
5High
6Critical
7Critical
8High
9High
10High
11Critical
12Critical
13High
14Critical
15High
16High
17Critical

Key Findings for Sentry

entelligence Strengths
  • Perfect precision (100%) - zero false positives
  • Catches Django ORM issues (negative slicing)
  • Identifies Celery serialization problems
  • Good at class definition time issues
entelligence Weaknesses
  • Lower recall than codex on this Python codebase
  • Missed abstract method implementation issues
  • Missed multiprocessing bugs
  • Missed Redis command compatibility issues

Context

How Other Tools Benchmark Code Review

Different AI code review tools use different benchmarking methodologies, which makes direct comparisons challenging. Here's how competitors approach evaluation:

Greptile's Approach

Greptile released their benchmark report using the same 5 repositories we tested (Cal.com, Sentry, Discourse, Keycloak, and Grafana). However, their methodology differs significantly:

  • Dataset size: Only one issue per PR in their evaluation (50 issues vs. our 67 bugs across 50 PRs)
  • Metrics: Reports only recall
  • Dataset availability: To our knowledge, Greptile did not release their ground truth dataset publicly

Takeaways

Key Takeaways

The test covered bugs that matter: race conditions, breaking changes, security vulnerabilities, and logic errors that crash production systems.

The gap came from understanding code relationships. When a function signature changes, when a cache eviction policy creates race conditions, or when authorization checks get skipped, these bugs require seeing beyond individual lines of code.

Performance across five languages stayed consistent. That consistency matters when teams work in polyglot codebases.

Summary

3. Conclusion

entelligence achieves the best F1 score (47.2%) among all evaluated tools by balancing recall and precision effectively.

It excels at TypeScript/JavaScript codebases (68.1% F1 on Cal.com), Schema validation errors, Async/await patterns, Authorization logic bugs, and SQL injection vulnerabilities.

However, it struggles with some concurrency and race patterns under load, Ruby-specific patterns, Abstract method implementations, and CSS/UI regressions.

The combination of entelligence with a tool that has higher recall on specific patterns (like Greptile for Java/Keycloak) could provide comprehensive coverage.

The 2026 AI Code Review BenchmarkThe 2026 AI Code Review BenchmarkThe 2026 AI Code Review Benchmark

The memory starts the day you connect a repo.

From that day on, every session, review and incident lands in the same index.

  • No credit card required
  • No training on customer code
  • Deploy in your cloud, your own keys