Turn AI spend into engineering ROISee it in action →

Bespoke Grades the World's AI. Entelligence Grades Bespoke's Code.

How Bespoke Labs uses Entelligence : Two of every three comments Entelligence leaves on a Bespoke PR get fixed before merge.

Bespoke Labs builds the environments that train reliable AI agents. Their open reasoning dataset OpenThoughts has racked up hundreds of thousands of downloads, their GEPA prompt-optimization framework runs across 200+ teams, and Terminal-Bench, their environment-based benchmark, is used by Anthropic, OpenAI, and Google DeepMind to evaluate agentic systems.

That last fact is the whole point. When the code you ship trains and grades other people's AI, a bad diff doesn't just break a feature. It can silently corrupt a training signal or skew a benchmark score that downstream labs are already citing. The bug never announces itself. It just quietly becomes bad data that everyone inherits.

So Bespoke holds its codebase to the same bar as its research. That's where Entelligence comes in.

The problem with shipping fast on ground truth

Bespoke runs a 50-person org across four active repos: a Postgres-backed product backend, a Python and TypeScript rubric-review pipeline, and the infra and deploy tooling underneath the code that trains and grades everyone else's agents. The team ships fast enough that no human can manually re-check every migration, every SQL query, and every auth check on every PR. On most teams, that's a velocity problem. Here, one missed check can poison the exact thing the rest of the industry trusts Bespoke to get right.

Three months, four repos, one reviewer that keeps up

Over three months, Entelligence's review covered 1,000+ PRs, ran 3000+ review executions, and left 5,000+ inline comments, averaging about four minutes per review.

The catches that would have shipped bad data

Each of these was flagged and fixed before it merged. Every one is the kind of silent fault a fast merge cadence lets through.

Findings


An authorization bypass on task reassignment

When a task lookup returned null, the membership check was skipped entirely, letting a pod lead reassign a task to any worker with no restriction. Flagged, fixed.

A stored XSS in the rollout message view

Code was interpolated straight into an HTML template literal without encoding, so a crafted payload could break out of the tag before the sanitizer ever ran. Flagged, fixed.

A migration that would never run

A new SQL migration file existed but was never registered with the runner, so it would silently skip in any fresh environment. Flagged, fixed.

A CI restore that reported success on failure

A piped database restore masked its real exit code because pipefail was never set, so a failed restore would still report green. Flagged, fixed.

A NaN injected into a SQL interval

An unvalidated day parameter could resolve to NaN, get concatenated into a Postgres interval string, and throw at query time. Flagged, fixed.

What the data shows

  • 1,600+ PRs reviewed, at about four minutes each

  • roughly 70% of review comments verified as fixes

  • about two-thirds of high-severity findings fixed before merge

  • a 99% review success rate

  • under half a percent of PRs reverted over the window

That's roughly two out of every three comments Entelligence left, acted on before merge, on a team shipping fast enough to need that many comments in the first place. Most review bots earn the opposite reputation, noise that developers learn to scroll past. Bespoke's engineers fixed the majority of what Entelligence flagged, and 99% reliability meant the reviews were there every time a PR was.

For a lab whose benchmarks the rest of the industry runs on, that's the difference that matters. Entelligence is the reviewer catching the silent faults before they become everyone's bad data.

We raised $5M to run your Engineering team on Autopilot

We raised $5M to run your Engineering team on Autopilot

Watch our launch video

Talk to Sales

Production reliability, solved.

The AI engineer that reviews every PR against your incident history, watches production, and self-heals when things break. The same class of bug will not ship twice.

Talk to Sales

Production reliability, solved.

Connect with our team to see how Entelliegnce helps engineering leaders with full visibility into sprint performance, Team insights & Product Delivery

Try Entelligence now