Skip to content
WHITEPAPER // 2026

Self-Improving Coding Agents: How Software Factories Learn From Their Own Mistakes

Models get better every quarter, but no model will ever ship knowing your codebase's contracts, conventions and incidents. That knowledge has to be learned from your own factory's mistakes, and stored somewhere the next session can use it.

· Aiswarya Sankar

A session replay. As an agent session plays, every file, MCP call and subagent it touches joins its context tree, with tokens and cost over time underneath. Reading traces like this is where self-improvement starts.

Why software factories need memory

A software factory is what engineering becomes when agents write most of the code. Humans set intent and review; agents plan, explore, edit and open PRs, hundreds of sessions a day. Throughput scales with the number of agents. Learning does not.

A human engineer who gets a review comment internalizes it and doesn't make that mistake again. An agent session starts cold every time. So the factory ships the same defect over and over, and the humans reviewing its output become its memory, which is the most expensive memory there is.

A self-improving factory closes that gap with a feedback loop, the way manufacturing did with statistical process control and machine learning did by training on its own errors. The loop has four stages:

  1. Pinpoint. Take a failure signal (a review comment, a failed test, an incident) and trace it back to the exact step in the agent session where it entered.
  2. Aggregate. Cluster those root causes across many sessions until isolated mistakes become recurring patterns.
  3. Inject. Turn each pattern into a learning, and give it to future sessions at the right moment, through the right mechanism.
  4. Prove. Measure against a control whether the learning actually prevents the error, and retire it if it doesn't.
Every review comment feeds a loop that makes the next session betterkeep what works, retire what doesn’t; the next review feeds the loop1 · Pinpointreview comment toturn, edit, ghost file2 · Aggregatecluster root causesinto learnings3 · InjectCLI adds context:rules/skill/MCP/hook4 · Prove80/20 holdout,per-learning scorecardWorld Modelsessions · PRs · issues · patterns · themes · learnings · people
The learning loop · 4 stages over one shared memory

Each stage is easy to state and hard to get right. Sections 1 to 4 cover the principles: why agents repeat mistakes, what makes a good error signal, how to read an agent trace, and why most failures turn out to be context failures. Sections 5 to 9 show a working implementation, the Entelligence Learning Loop, with numbers from an eight-week rollout. Sections 10 and 11 look at where the loop goes next.

The core idea is simple. Models get better every quarter, but no model will ever ship knowing your codebase's contracts, conventions and incidents. That knowledge has to be learned from your own factory's mistakes, and stored somewhere the next session can use it.

1. The problem: agents repeat themselves

Anyone who reviews agent-written PRs at volume sees this. Cluster a few weeks of substantive review comments and the same failures come back week after week, across different engineers, different repos and different agents.

The pattern is consistent. An agent passes user_id where the attribution contract expects cohort_id. It hand-rolls pagination offsets in a repo that has a paginate() helper. It calls db.session.execute_raw, a SQLAlchemy-shaped method that doesn't exist in the in-house ORM wrapper. A reviewer catches it, the engineer fixes it, the PR merges, and three days later a different session makes the identical mistake.

The reason is structural. Every agent session starts cold. A human engineer who gets that review comment once never makes the mistake again, because they remember. The agent has no equivalent of that memory. Static rules files (CLAUDE.md, AGENTS.md, .cursorrules) are the industry's current answer, but they are hand-written, rarely updated, never measured, and grow until agents stop reading them carefully.

In the rollout described later in this paper, 94% of substantive review comments across six repos traced back to a specific agent edit. These aren't human mistakes slipping through. They are agent mistakes, and they are learnable.

2. Code review is the error sensor

Every learning loop needs a reliable signal of failure. For coding agents, the best signal already exists: the code review. A review comment is a labeled example, written by an expert, pinned to a file and line, attached to a PR that an agent session produced.

To feed a learning loop, a review signal needs three properties that ordinary code review doesn't guarantee:

  • Classified. Substantive issues (a broken contract, a wrong API, a violated convention, security, logic) are separated from nits. Nits stay out of the loop by default, because training agents against style preferences produces noise, not quality.
  • Linked to the edit that caused it. A comment on views.py L140 has to join to the exact edit event in the agent trace (say, edit #612) that wrote that line. Without that link you know what's wrong, but not where the agent went wrong.
  • Inclusive of human reviewers. Comments from engineers, not just automated reviewers, flow into the same pipeline, so the loop learns the team's real standards.

Precision matters more than recall. A loop fed false positives teaches agents to avoid problems that don't exist, and spends context doing it. A factory's memory can only be as good as its failure signal.

3. Trace forensics: down to the step and turn

A review comment tells you what went wrong. The agent trace tells you how and why. The first step is to normalize sessions from every harness in use (Claude Code, Codex, Cursor and others) into a single event model.

Every event in a session is typed and indexed:

Event typeWhat we captureWhy it matters
PromptUser instructions and follow-upsDid the human give the agent the constraint?
ExplorationFile reads, greps, directory listingsWhat did the agent actually look at before editing?
EditFile, line range, diff, event indexThe line a reviewer later flags
SubagentSpawn, calls, cost, handoff summaryWhat was learned in a subagent and lost on return
MCP / toolTool calls, arguments, resultsWhich context sources were queried, and what came back
ErrorFailed tool calls, permission denials, test failuresWhere the agent hit friction and improvised
SkillSkills loaded and whenWhether available guidance was ever used

The example session in this paper is prod-db agent PR review investigation plan, run on Claude Code for 4h 45m, 921 events, 8 subagents, 11 errors, $13.11. It produced PR #4182, which drew three substantive review comments and one nit.

For each comment we compute a causal path: the minimal chain of events that led from what the agent knew to the flagged edit. For the attribution bug, that path is four events long:

  1. #388, Sub 4 (Explore): reads model_attribution.py, which defines the contract.
  2. #402, Sub 4: returns a handoff summary to the main agent that omits cohort_id.
  3. #598, Main: reads views.py.
  4. #612, Main: edits views.py L140, passing user_id. Flagged in review.

The subagent had the right answer. The main agent never received it. You cannot see that from the diff; you can only see it from the trace.

We also compute ghost nodes: files that the fix required but the agent never opened. In this case attribution_serializer.py defines the contract and was never read. Ghosts are often the single most actionable output of the analysis, because they name exactly what context was missing.

4. Root cause: most failures are context failures

The model usually isn't the problem. When we classify causal paths, the large majority of substantive errors trace to the agent not having, not finding, or not keeping the right context at the moment it edited. Only a minority are genuine reasoning failures.

We classify every causal path into one of six root causes:

Root causeWhat happened in the traceExampleComments / session before learnings
Missing contextThe fact the agent needed wasn't in anything it readFE charts redeclare cohort types that exist in @/types/cohort0.38
Ignored conventionThe convention exists in the repo, but the agent never opened itRead specs/, never CONVENTIONS.md; hand-rolled pagination0.31
Lost subagent handoffA subagent found the answer; the summary back to main dropped itcohort_id contract read at #388, lost at #4020.24
Hallucinated APIThe agent used a plausible method from public training dataexecute_raw exists in SQLAlchemy, not in the house wrapper0.19
Plan driftThe plan was right; execution diverged from it over a long sessionEdits late in a 900-event session contradict the plan file0.17
Tool / permissionA tool failed or was denied and the agent worked around it badlyPermission denied, then an Edit bypass that skipped a hook0.13

Three of these (missing context, ignored convention, hallucinated API) are fixed by supplying a fact. One (lost handoff) is fixed by changing how the harness moves context between agents. Plan drift and tool failures need timing: the right reminder, at the right step.

That split is the core design insight behind the memory layer. "Add more to CLAUDE.md" solves only the first category, and poorly. Each root cause needs a different delivery mechanism, which is why Section 7 treats where, when and how context is added as an experimental variable, not a constant.

From principles to practice: the Entelligence Learning Loop

The four stages above are a design pattern; any team can build toward them. At Entelligence we've built the full loop as a product, and the rest of this paper shows how it works and what it measured.

We started from the error sensor. Entelligence's code reviewer ranked #1 of 8 tools on real production bugs (47.2% F1) in our Code Review Benchmark 2026, which gave us a high-precision failure signal to learn from. From there we added trace ingestion, root-cause analysis, a shared memory we call the World Model, and the Entelligence CLI, which builds each agent session with the right learnings installed. As far as we know, no one else closes this loop end to end.

The Learning Loop has one view per stage: Pinpoint (one session), Aggregate (all sessions), Inject (the audit), and Prove (over time). Sections 5 to 9 walk through each. The numbers come from an eight-week rollout, Aug 4 to Sep 29, across 1,284 Claude Code sessions in six repos. The headline result: sessions built with learnings drew 53% fewer substantive review comments than a holdout, at $2.41 of injected context in total.

5. Where in the session each error was made, and why

For every flagged line, we show the engineer the exact moment in the session the error entered, the path that led there, and what the agent should have read. The session replay renders each review comment on the session timeline at the edit that wrote it, then draws the causal path back through subagents and reads.

Pinpoint view: PR #4182 session with review comments mapped onto the context tree and causal path
Stage 1, Pinpoint. The PR #4182 session in the Learning Loop. Each review comment sits on the context tree at the file it flagged, with the path to the flagged edit in red and the ghost file dashed. The right panel shows the cause, the causal path event by event, and the learning it was clustered into.

The diagram below pulls out that one causal path so it can be read at a glance.

The contract was read at #388 and lost at #402, so #612 broke itprod-db agent PR review investigation plan · Claude Code · 921 events · 8 subagentsSub 4 · Explore subagentMain agentReviewflagged#388 · Readmodel_attribution.pyhas the contract#402 · Handoffsummary back to mainomits cohort_idGhost: attribution_serializer.pydefines the contract; never opened before #612the missing context, named#598 · Readviews.pyno contract in context#612 · Editviews.py L140passes user_idSame PR, same analysis for every comment#640 views.py L188: hand-rolled pagination · ignored convention#701 controller.py L52: execute_raw · hallucinated APIReview comment@mkaur on PR #4182breaks the contractwhere the error entered the sessionghost: needed by the fix, never read
Session replay · causal path for one review comment

The subagent did its job; the handoff didn't. That distinction changes the fix: this isn't a rule to add to CLAUDE.md, it's a learning about how Explore subagents must report back (L-9), plus a path-specific fact about the attribution contract (L-14).

In the product, the replay also shows the session's full context tree: every file, MCP call, subagent and insight source the agent touched, sized by how often, with ghost nodes drawn dashed. Across the whole PR #4182 session, that tree covered 328 exploration events, 113 tool calls, 11 MCP calls and 11 errors.

6. From one session to recurring patterns

One causal path is an anecdote. The same causal path in five sessions is a pattern worth teaching. We cluster root-caused comments across every session in the org, by root cause, trigger path and the contract or convention involved, and propose a learning when a cluster crosses an evidence threshold.

Aggregate view: learnings list with L-14 selected, its trigger, delivery, root cause and evidence
Stage 2, Aggregate. All learnings across 1,284 sessions. L-14 is selected: its trigger, delivery mechanism, root cause, the four PR comments behind it (each with a replay), and where it sits in the approval lifecycle.
Four review comments in two repos became one learning, L-14Clustered by root cause: missing context + lost subagent handoffPR #4182 · prod-db · Sep 27views.py L140: expects cohort_idPR #4120 · data-pipeline · Sep 22backfill.py L88: serializer bypassedPR #4077 · prod-db · Sep 18views.py L212: user_id in payloadPR #4012 · prod-db · Sep 12reports.py L40: serializer skippedL-14 · ActiveAttribution changes in viewsgo through AttributionSerializer.model_attribution expectscohort_id, not user_id.9 comments · 5 sessions · 2 reposDeliveryWorld Model MCPtrigger:backend/**/views.py+ diff touchesattribution
Aggregation · review evidence clustered into one learning

Each learning carries its evidence trail: every PR comment it came from, with a link and a replay of the session that produced it. That trail is what a human reviews at the approval gate, and what makes each learning auditable later.

Across 1,284 sessions and 412 substantive comments over eight weeks, the loop produced six learnings:

LearningRuleRoot causeEvidenceDeliveryStatus
L-14Attribution in views goes through AttributionSerializer; expects cohort_idLost handoff + missing context9 comments · 5 sessions · 2 reposWorld Model MCPActive
L-11ORM wrapper has no execute_raw; use db.run_sql(query, params)Hallucinated API5 comments · 5 sessions · 2 reposPre-edit hookActive
L-12FE charts import cohort types from @/types/cohortMissing context8 comments · 6 sessions · 1 repoRules fileActive
L-9Explore subagents return every file path and contract they readLost handoff7 comments · 6 sessions · 3 reposSkillApproved
L-15List endpoints use core.pagination.paginate()Ignored convention6 comments · 4 sessions · 1 repoRules fileProposed
L-6Name DataFrame variables with a _df suffixStyle preference11 nits · 3 sessions · 1 repoRules fileRetired

Note L-9: it spans three repos because it isn't about any codebase. It's about the harness itself. Some of the most valuable learnings fix how agents work, not what they know.

7. Closing the loop: injection through the Entelligence CLI

A learning is worthless if it doesn't reach the next session at the moment it matters. The Entelligence CLI builds each agent session: before the harness starts, it resolves which learnings apply to this repo, this task and these paths, and installs them through one of four mechanisms.

MechanismHow it's deliveredBest forCost profile
Rules fileCompiled into CLAUDE.md / AGENTS.md at buildShort, always-true conventions (L-12: import cohort types from @/types/cohort)Always in context; cheap per item, expensive in aggregate
SkillBundled as a skill, loaded when the task type matchesProcedural fixes (L-9: Explore subagents must return every contract they read)Zero until the task matches
World Model MCPRegistered at build; retrieved when trigger paths matchRich, path-specific knowledge (L-14: attribution goes through AttributionSerializer)Pay only on retrieval, ~1.1K tokens
Pre-edit hookRuns deterministically before edits to trigger pathsHard constraints that must never be violated (L-11: no execute_raw, use db.run_sql)~0.3K tokens; enforced, not suggested

Every learning carries a trigger, a path glob plus a diff condition (for example backend/**/views.py · diff touches attribution). Triggers are what keep context lean: across 1,027 sessions built with learnings, the CLI made 2,356 insertions at an average of just +0.9K tokens per session.

Inject view: injection audit of which learnings the CLI added to each session and whether they were used
Stage 3, Inject. The injection audit. For every session the CLI built: which learnings it added and through which mechanism, the tokens added, whether the agent actually used them in the trace, and what review said. Holdout sessions are marked, and a learning that was retrieved but ignored is flagged in red.

What we experiment on

We treat injection as a four-variable experiment, and run it continuously:

  • What to add: the rule itself, the evidence behind it, a pointer to the file that defines it, or all three. Pointers to ghost files often outperform prose rules.
  • Where to add it: rules file, skill, MCP or hook. The same learning can succeed as a hook and fail as a rules line.
  • When it fires: at session build, at task-type match, at first read of a trigger path, or immediately before an edit. Later is usually better; the agent acts on what's in its recent context.
  • How much: per-session token budget, ranking when many learnings match, and suppression of learnings that have stopped paying off.

The audit log keeps us honest about delivery versus use. A learning can be delivered and still ignored. In PR #4214, L-14 was retrieved, then ignored, and the error recurred. In PR #4182, the L-11 hook never fired because the agent wrote execute_raw through an Edit path that bypassed it. Both are logged as failures of the mechanism, not successes of delivery, and both feed back into the next round of experiments.

Governance

Learnings move through an explicit lifecycle: Proposed (clustered from review comments) → Approved (human gate; nothing enters a session without approval) → Active (added by the CLI) → Retired (removed when it stops paying off). Engineers stay in control of what their agents are told.

8. The World Model: where the memory lives

The World Model is a graph of everything the loop has learned about how agents behave in your org. It stores every session, PR, review comment and person as nodes, and the causal links between them as edges. In the view below, the top 238 patterns, people and themes are drawn from a graph of 162,130 nodes.

World model view in Entelligence: issues, themes, sessions, patterns, PRs and people on a globe

In this loaded view, the graph holds:

Node typeCount in viewWhat it represents
Issue429A specific failure, e.g. "Stale API Key Retried on 401", "Localhost Fallback in Prod"
Theme150A cluster of related work or failures, e.g. "Ellie CLI audit and merge"
Session132An agent trace, with cost, e.g. "claude-code · $376"
Pattern69A recurring root cause across sessions, e.g. a lost subagent handoff
PR40The diff a session produced, e.g. PR #5220
Person19The engineers whose sessions feed the graph

The view has three lenses: Recurring Issues (what keeps breaking), Coding Agent (how agents behave), and Merged (both, linked). 283 links connect the loaded nodes, and hovering an orb traces them. Ask Ellie answers questions against the same graph in natural language.

Every learning traces back to the edit and the session behind itWorld Model schema: node types and the links between themrunsproducesreceivesclusters intobecomesrolls up intoadded at build by the CLIPersonengineer who ranthe sessionSessionfull agent trace:every turn + tool callPRthe diff thesession producedIssuea review comment,linked to its edit #Learningrule + trigger +delivery; approvedPatternrecurring root causeacross sessionsThemea cluster of patterns,e.g. auth, attribution
World Model schema · 7 node types

The schema is what makes the memory trustworthy. Every learning resolves back through a pattern to specific issues, to the edits that caused them, to the sessions and people involved. Nothing an agent is told is unsourced.

Agents read the graph live through the world_model.search MCP tool, which the CLI registers at build. That is the delivery path for rich, path-specific learnings like L-14: rather than bloating every session's rules file, the agent retrieves the relevant slice of memory only when it touches backend/**/views.py.

9. Proving it: the same errors, measured over time

Sessions built with learnings drew 53% fewer substantive review comments than the holdout (0.62 vs 1.31 per session; 95% CI −41% to −63%). Every claim of improvement in this system is made against a control, not against last month.

Prove view: comments and review rounds vs holdout, comments by root cause, per-learning scorecard, holdout and cost
Stage 4, Prove. The impact view: substantive comments and review rounds against the holdout, comments by root cause before and after, the per-learning scorecard, and the cost of injection. The charts and tables below break these out.

How we measure. Since Aug 4, 20% of sessions are a standing holdout. For those sessions, the CLI still matches triggers and logs which learnings would have been added, but adds nothing. That gives us a like-for-like control: same repos, same engineers, same agents, same weeks. The holdout itself improved about 10% over the period, which is the team-effect baseline (people get better, models get better). We report the gap above that line, not the raw drop.

With learnings, comments per session fell to 0.62; holdout stayed at 1.31Substantive review comments per session, by ISO week, 2026 · 80/20 split0.00.51.01.5W31W32W33W34W35W36W37W38learnings switched onWith learnings, W31: 1.42 comments / sessionWith learnings, W32: 1.38 comments / sessionWith learnings, W33: 1.21 comments / sessionWith learnings, W34: 1.05 comments / sessionWith learnings, W35: 0.88 comments / sessionWith learnings, W36: 0.74 comments / sessionWith learnings, W37: 0.66 comments / sessionWith learnings, W38: 0.62 comments / sessionWith learnings 0.62Holdout, W31: 1.45 comments / sessionHoldout, W32: 1.40 comments / sessionHoldout, W33: 1.37 comments / sessionHoldout, W34: 1.35 comments / sessionHoldout, W35: 1.33 comments / sessionHoldout, W36: 1.30 comments / sessionHoldout, W37: 1.32 comments / sessionHoldout, W38: 1.31 comments / sessionHoldout 1.31
Learning Loop review · 1,284 sessions · W31–W38 2026

The lines separate the week learnings switch on and keep separating as more learnings go live. The holdout barely moves, which is the point: the improvement comes from the memory, not from the model or the calendar.

Because every comment is root-caused, we can also see which kinds of failure the memory fixes.

Handoff and hallucinated-API errors fell most; plan drift fell leastSubstantive comments per session by root cause, sessions with learningsBefore learnings, W29–W32After, W35–W38change0.00.10.20.30.4Missing contextMissing context, before: 0.38 / sessionMissing context, after: 0.16 / session−58%Ignored conventionIgnored convention, before: 0.31 / sessionIgnored convention, after: 0.14 / session−55%Lost subagent handoffLost subagent handoff, before: 0.24 / sessionLost subagent handoff, after: 0.08 / session−67%Hallucinated APIHallucinated API, before: 0.19 / sessionHallucinated API, after: 0.05 / session−74%Plan driftPlan drift, before: 0.17 / sessionPlan drift, after: 0.13 / session−24%Tool / permissionTool / permission, before: 0.13 / sessionTool / permission, after: 0.06 / session−54%
Learning Loop review · sessions with learnings · 4 weeks before vs 4 weeks after

This is the honest part of the result, and the one we find most useful. Context injection crushes failures caused by missing facts and lost handoffs. It does much less for plan drift, which is a long-horizon execution problem. That tells us where to push next: mid-session reminders timed to plan checkpoints, rather than more context at build.

MetricWith learningsHoldout
Sessions1,027257
Review rounds to merge1.31.9
Agent errors per session4.16.8

Per learning, not just in aggregate. Each learning gets its own scorecard, so a weak learning can't hide behind a strong one:

LearningDeliveryTriggeredUsedRecurredPreventedCost / sessionVerdict
L-11Pre-edit hook6458314$0.001Keep
L-12Rules file413348$0.001Keep
L-14World Model MCP383529$0.002Keep
L-6Rules file4426181$0.001Retired

L-6 is the loop working as designed. It was a style rule learned from nits; it recurred 18 of 44 times, prevented almost nothing, and was retired. Memory that doesn't pay for itself gets removed.

Cost. Injection averages +0.9K tokens and $0.0023 per session: $2.41 across eight weeks. L-14 alone costs $0.002 per session and prevents about one review round-trip every eight sessions. That's the economic case in one line: a fraction of a cent of context replaces an engineer's review cycle.

10. Beyond code review: every signal your org already has

Code review is the first sensor, not the only one. The same loop works for any signal that says "an agent-written change caused a problem", and on the upgraded Entelligence plan these plug into the World Model today:

SourceWhat it teaches the agentExample learning
Production incidents (PagerDuty, incident.io)Which changes broke prod, and the fix that resolved them"Migrations on large tables must be batched; an unbatched one locked the table in prod"
Alerts and errors (Sentry, Datadog)Runtime failures traced back to the PR and session that shipped them"This endpoint's null team_model crashes in prod; guard it"
Designs and specs (Figma, Notion, Linear)Intended behavior and decisions the code doesn't show"Checkout copy and states follow the approved Figma flow"
Customer tickets (Zendesk, Intercom, Slack)Real user pain linked to the code that caused it"Large exports time out for enterprise customers; stream them"

Every source is normalized the same way: an event, linked to code, linked to the session that wrote it, clustered into a learning, gated by a human, delivered by the CLI, and measured by holdout. The loop is source-agnostic by design.

This is the long-term shape of the product: a production-aware memory for every coding agent in the org, so that the incident your on-call engineer fixed at 3 a.m. is something no agent in your company will ever cause again.

11. Open source, and where this goes

The world model visualization is open source: Entelligence-AI/agentplay. Point it at your own agent sessions to see the patterns, people and themes in your graph, and how they connect.

We're open-sourcing the view because we want the whole industry looking at agent traces this way. Agents will keep getting smarter. What they won't get on their own is the memory of your codebase: the contracts, conventions, incidents and decisions that make your code yours. We are building the layer that captures that memory, delivers it, and proves it works. As far as we know, no one else closes this loop end to end, and we intend to define what it looks like.

Talk to Sales

Production reliability, solved.

The AI engineer that reviews every PR against your incident history, watches production, and self-heals when things break. The same class of bug will not ship twice.