Benchmark report
Entelligence Code Review Benchmark 2026
Entelligence was benchmarked against 7 leading AI code review tools (CodeRabbit, Greptile, Copilot, Graphite, Codex, Claude, and Bugbot) on real production bugs.
- Reviewers
- 8
- Production bugs
- 67
- Repositories
- 5
Overview
Introduction
AI code review tools promise to catch bugs before production. But do they actually work?
We tested 8 AI code reviewers in total: Entelligence, CodeRabbit, Greptile, Copilot, Graphite, Claude, Codex, and Bugbot on 67 real production bugs from open-source projects. Each bug shipped to production broke something and required a fix in a later commit.
The bugs included race conditions, security vulnerabilities, breaking API changes, and logic errors across five different repositories: Cal.com (TypeScript), Sentry (Python), Discourse (Ruby), Keycloak (Java), and Grafana (Go).
Results
Results (F1 scores)
Full results table
| Rank | Reviewer | F1 Score | Found/Golden |
|---|---|---|---|
| 1 | Entelligence | 47.2% | 30/67 |
| 2 | Codex | 45.4% | 27/67 |
| 3 | Claude | 42.8% | 29/67 |
| 4 | Bugbot | 39.4% | 31/67 |
| 5 | Greptile | 36.9% | 31/67 |
| 6 | CodeRabbit | 33.0% | 32/65 |
| 7 | Copilot | 22.6% | 23/67 |
| 8 | Graphite | 13.4% | 5/67 |
Evidence
See The Proof
How we tested
Methodology
Building a Golden Dataset for Code Review Evaluation
Benchmarking AI code reviewers requires a ground truth: a set of golden comments representing what a thorough, context-aware review should look like. Catch rate alone - whether a tool flags a known bug—doesn't capture comment quality, relevance, or actionability. We set out to build a golden dataset that could measure these dimensions.
An Ensemble Approach to Ground Truth
Rather than relying on a single model or a single human's judgment, we built our golden dataset using an ensemble of frontier LLMs: Claude Opus, Gemini 2.5, and GPT-5. Each model received identical context for every pull request diff - including the inbound and outbound dependencies of the changed code. We extracted these dependencies using Language Server Protocol (LSP), which allowed us to programmatically trace function calls, imports, and type relationships across the codebase. This dependency context is critical; it's what separates superficial line-by-line feedback from reviews that understand how a change ripples through a codebase.
We prompted each model independently to generate review comments. This gave us three distinct perspectives on every diff, each shaped by the model's own reasoning patterns and training.
Consensus Through Voting
Three models means three (often different) sets of comments. To distill these into a single golden dataset, we implemented a majority voting process. A separate LLM call - GPT-4o - acted as the arbiter, evaluating which comments across the three sources addressed the same issues and selecting those with majority agreement. Comments that only one model flagged were discarded; comments that two or three models independently identified became candidates for the golden set.
This voting mechanism serves two purposes: it filters out model-specific hallucinations or stylistic quirks, and it surfaces issues that multiple sophisticated models agree are worth flagging. If Opus, Gemini, and GPT-5 all independently identify the same problem, that's a strong signal it belongs in the ground truth.
Human Validation
LLM consensus isn't infallible. To ensure quality, reviewers at our company manually examined the voting-selected comments. This human validation step caught edge cases where models agreed on something incorrect or where the voting process produced ambiguous results. The final golden dataset represents the intersection of multi-model consensus and human expert judgment.
Through this process, we generated 67 golden comments. We used the same repositories featured in Greptile's benchmark - Sentry, Cal.com, Grafana, Keycloak, and Discourse - so our results are directly comparable to existing industry benchmarks. The full dataset, including all golden comments, is available in our repository.
Evaluation Methodology
With our golden dataset established, we ran the benchmark. We cloned each repository and opened pull requests mirroring the original bug-introducing commits. We triggered each AI code review tool (Codex, Greptile, CodeRabbit, Entelligence, Graphite, Copilot, Claude Code, and Bugbot) on these PRs under default settings, excluding dependency bot PRs to focus on meaningful code changes.
For each tool, we collected every review comment and compared it against our golden comments using Claude Sonnet 4.5 as a semantic similarity judge. The LLM evaluated whether each AI-generated comment captured the same issue and intent as the corresponding golden comment.
Analysis
Combined Evaluation Analysis
1. Executive Summary
Aggregate Metrics (All Repositories)
| Rank | Reviewer | F1 Score | Recall | Precision | Found/Golden |
|---|---|---|---|---|---|
| 1 | Entelligence | 47.2% | 44.8% | 50.0% | 30/67 |
| 2 | Codex | 45.4% | 40.3% | 51.9% | 27/67 |
| 3 | Claude | 42.8% | 43.3% | 42.3% | 29/67 |
| 4 | Bugbot | 39.4% | 46.3% | 34.4% | 31/67 |
| 5 | Greptile | 36.9% | 46.3% | 30.7% | 31/67 |
| 6 | CodeRabbit | 33.0% | 49.2% | 24.8% | 32/65 |
| 7 | Copilot | 22.6% | 34.3% | 16.8% | 23/67 |
| 8 | Graphite | 13.4% | 7.5% | 66.7% | 5/67 |
* Note: CodeRabbit was evaluated on 65 total PRs (14 out of 16 in Cal.com) due to PR/dataset availability during testing. All other tools were evaluated on the full 67 PRs.
2. Per-Repository Performance
entelligence Performance by Repository
| Rank | Repository | F1 | Recall | Precision | Found/Golden |
|---|---|---|---|---|---|
| 1 | Cal.com | 68.1% | 81.2% | 58.6% | 13/16 |
| 2 | Grafana | 50.0% | 33.3% | 100.0% | 3/9 |
| 3 | Sentry | 45.5% | 29.4% | 100.0% | 5/17 |
| 4 | Keycloak | 38.5% | 33.3% | 45.5% | 4/12 |
| 5 | Discourse | 30.3% | 38.5% | 25.0% | 5/13 |
TypeScript
Cal.com - Issue Detection Matrix
Summary Metrics
| Rank | Reviewer | F1 Score | Recall | Precision | Found/Golden |
|---|---|---|---|---|---|
| 1 | Entelligence | 68.1% | 81.2% | 58.6% | 13/16 |
| 2 | Codex | 58.8% | 56.2% | 61.5% | 9/16 |
| 3 | Claude | 52.2% | 50.0% | 54.5% | 8/16 |
| 4 | Bugbot | 50.0% | 68.8% | 39.3% | 11/16 |
| 5 | Greptile | 47.4% | 56.2% | 40.9% | 9/16 |
| 6 | CodeRabbit | 46.0% | 78.6% | 32.6% | 11/14 |
| 7 | Copilot | 37.0% | 62.5% | 26.3% | 10/16 |
| 8 | Graphite | 0.0% | 0.0% | 0.0% | 0/16 |
* Note: CodeRabbit was evaluated on 65 total PRs (14 out of 16 in Cal.com) due to PR/dataset availability during testing. All other tools were evaluated on the full 67 PRs.
Per-PR Issue Detection Matrix
* The ↗ symbol indicates the tool successfully scanned and commented on the PR. It does not indicate whether the tool successfully found the specific golden bug. Please refer to the summary tables for exact detection rates.
Cal.com Issues (16 total)
Tip: Click any ✓/✗ to view the PR evidence.
Key Findings for Cal.com
- Strong critical issue detection (5/6 = 83%)
- Catches complex OAuth token handling issues
- Identifies async/await problems (forEach with async)
- Good at null/undefined access patterns
- Generates more false positives than codex (12 vs 5)
- Missed timezone-related issue (organizerTimeZone)
- Some false positives are valid issues not in golden set
Ruby
Discourse – Issue Detection Matrix
13 Golden Issues
Summary Metrics
| Rank | Reviewer | F1 Score | Recall | Precision | Found/Golden |
|---|---|---|---|---|---|
| 1 | Bugbot | 44.4% | 46.2% | 42.9% | 6/13 |
| 2 | Codex | 38.1% | 30.8% | 50.0% | 4/13 |
| 3 | Greptile | 35.7% | 46.2% | 29.2% | 6/13 |
| 4 | Entelligence | 30.3% | 38.5% | 25.0% | 5/13 |
| 5 | Claude | 26.4% | 30.8% | 23.1% | 4/13 |
| 6 | CodeRabbit | 25.9% | 53.8% | 17.0% | 7/13 |
| 7 | Graphite | 25.0% | 15.4% | 66.7% | 2/13 |
| 8 | Copilot | 11.0% | 15.4% | 8.6% | 2/13 |
Per-PR Issue Detection Matrix
* The ↗ symbol indicates the tool successfully scanned and commented on the PR. It does not indicate whether the tool successfully found the specific golden bug. Please refer to the summary tables for exact detection rates.
Discourse Issues (13 total)
Tip: Click any ✓/✗ to view the PR evidence.
Key Findings for Discourse
- Catches SQL injection vulnerability (security-critical)
- Identifies null pointer exceptions
- Good at localization/i18n issues
- Missed Ruby syntax errors
- Lower recall than CodeRabbit and Greptile on this codebase (Ruby)
- Misses some nil dereference patterns
- CSS/theme-related issues completely missed
Go
Grafana Repository (9 Golden Issues)
Grafana – Issue Detection Matrix
Summary Metrics
| Rank | Reviewer | F1 Score | Recall | Precision | Found/Golden |
|---|---|---|---|---|---|
| 1 | Codex | 64.5% | 66.7% | 62.5% | 6/9 |
| 2 | Claude | 56.3% | 55.6% | 57.1% | 5/9 |
| 3 | Entelligence | 50.0% | 33.3% | 100.0% | 3/9 |
| 4 | CodeRabbit | 44.4% | 33.3% | 66.7% | 3/9 |
| 5 | Graphite | 44.4% | 33.3% | 66.7% | 3/9 |
| 6 | Bugbot | 33.8% | 44.4% | 27.3% | 4/9 |
| 7 | Greptile | 31.6% | 66.7% | 20.7% | 6/9 |
| 8 | Copilot | 23.5% | 33.3% | 18.2% | 3/9 |
Per-PR Issue Detection Matrix
* The ↗ symbol indicates the tool successfully scanned and commented on the PR. It does not indicate whether the tool successfully found the specific golden bug. Please refer to the summary tables for exact detection rates.
Grafana Issues (9 total)
Tip: Click any ✓/✗ to view the PR evidence.
Key Findings for Grafana
- Perfect precision (100%) - no false positives
- Catches initialization nil issues
- Good at identifying parameter mismatches
- Low recall (33.3%) on Go codebase
- Misses race condition issues (0/3)
- Misses stub function issues
Java
Keycloak Repository (12 Golden Issues)
Keycloak – Issue Detection Matrix
Summary Metrics
| Rank | Reviewer | F1 Score | Recall | Precision | Found/Golden |
|---|---|---|---|---|---|
| 1 | Greptile | 52.6% | 41.7% | 71.4% | 5/12 |
| 2 | Claude | 45.5% | 41.7% | 50.0% | 5/12 |
| 3 | Entelligence | 38.5% | 33.3% | 45.5% | 4/12 |
| 4 | Bugbot | 24.0% | 25.0% | 23.1% | 3/12 |
| 5 | CodeRabbit | 22.2% | 41.7% | 15.2% | 5/12 |
| 6 | Copilot | 18.7% | 25.0% | 15.0% | 3/12 |
| 7 | Codex | 18.2% | 16.7% | 20.0% | 2/12 |
| 8 | Graphite | 0.0% | 0.0% | 0.0% | 0/12 |
Per-PR Issue Detection Matrix
* The ↗ symbol indicates the tool successfully scanned and commented on the PR. It does not indicate whether the tool successfully found the specific golden bug. Please refer to the summary tables for exact detection rates.
Keycloak Issues (12 total)
Tip: Click any ✓/✗ to view the PR evidence.
Key Findings for Keycloak
- Catches HTML sanitizer issues (validation order, case sensitivity)
- Identifies authorization ID mismatches
- Reasonable precision (36.4%)
- Low critical issue detection (20%)
- Missed passkey authentication bugs
- Missed duplicate null check bug
Python
Sentry - Issue Detection Matrix
Summary Metrics
| Rank | Reviewer | F1 Score | Recall | Precision | Found/Golden |
|---|---|---|---|---|---|
| 1 | Codex | 46.2% | 35.3% | 66.7% | 6/17 |
| 2 | Entelligence | 45.5% | 29.4% | 100.0% | 5/17 |
| 3 | CodeRabbit | 37.5% | 35.3% | 40.0% | 6/17 |
| 4 | Bugbot | 35.0% | 41.2% | 30.4% | 7/17 |
| 5 | Claude | 34.3% | 41.2% | 29.4% | 7/17 |
| 6 | Greptile | 24.5% | 29.4% | 21.1% | 5/17 |
| 7 | Copilot | 19.7% | 29.4% | 14.8% | 5/17 |
| 8 | Graphite | 0.0% | 0.0% | 0.0% | 0/17 |
Per-PR Issue Detection Matrix
* The ↗ symbol indicates the tool successfully scanned and commented on the PR. It does not indicate whether the tool successfully found the specific golden bug. Please refer to the summary tables for exact detection rates.
Sentry Issues (17 total)
Tip: Click any ✓/✗ to view the PR evidence.
Key Findings for Sentry
- Perfect precision (100%) - zero false positives
- Catches Django ORM issues (negative slicing)
- Identifies Celery serialization problems
- Good at class definition time issues
- Lower recall than codex on this Python codebase
- Missed abstract method implementation issues
- Missed multiprocessing bugs
- Missed Redis command compatibility issues
Context
How Other Tools Benchmark Code Review
Different AI code review tools use different benchmarking methodologies, which makes direct comparisons challenging. Here's how competitors approach evaluation:
Greptile's Approach
Greptile released their benchmark report using the same 5 repositories we tested (Cal.com, Sentry, Discourse, Keycloak, and Grafana). However, their methodology differs significantly:
- Dataset size: Only one issue per PR in their evaluation (50 issues vs. our 67 bugs across 50 PRs)
- Metrics: Reports only recall
- Dataset availability: To our knowledge, Greptile did not release their ground truth dataset publicly
Takeaways
Key Takeaways
The test covered bugs that matter: race conditions, breaking changes, security vulnerabilities, and logic errors that crash production systems.
The gap came from understanding code relationships. When a function signature changes, when a cache eviction policy creates race conditions, or when authorization checks get skipped, these bugs require seeing beyond individual lines of code.
Performance across five languages stayed consistent. That consistency matters when teams work in polyglot codebases.
Summary
3. Conclusion
entelligence achieves the best F1 score (47.2%) among all evaluated tools by balancing recall and precision effectively.
It excels at TypeScript/JavaScript codebases (68.1% F1 on Cal.com), Schema validation errors, Async/await patterns, Authorization logic bugs, and SQL injection vulnerabilities.
However, it struggles with some concurrency and race patterns under load, Ruby-specific patterns, Abstract method implementations, and CSS/UI regressions.
The combination of entelligence with a tool that has higher recall on specific patterns (like Greptile for Java/Keycloak) could provide comprehensive coverage.
The memory starts the day you connect a repo.
From that day on, every session, review and incident lands in the same index.
- No credit card required
- No training on customer code
- Deploy in your cloud, your own keys





