Self-Improving Coding Agents: How Software Factories Learn From Their Own Mistakes
Models get better every quarter, but no model will ever ship knowing your codebase's contracts, conventions and incidents. That knowledge has to be learned from your own factory's mistakes, and stored somewhere the next session can use it.
· Aiswarya Sankar
Why software factories need memory
A software factory is what engineering becomes when agents write most of the code. Humans set intent and review; agents plan, explore, edit and open PRs, hundreds of sessions a day. Throughput scales with the number of agents. Learning does not.
A human engineer who gets a review comment internalizes it and doesn't make that mistake again. An agent session starts cold every time. So the factory ships the same defect over and over, and the humans reviewing its output become its memory, which is the most expensive memory there is.
A self-improving factory closes that gap with a feedback loop, the way manufacturing did with statistical process control and machine learning did by training on its own errors. The loop has four stages:
- Pinpoint. Take a failure signal (a review comment, a failed test, an incident) and trace it back to the exact step in the agent session where it entered.
- Aggregate. Cluster those root causes across many sessions until isolated mistakes become recurring patterns.
- Inject. Turn each pattern into a learning, and give it to future sessions at the right moment, through the right mechanism.
- Prove. Measure against a control whether the learning actually prevents the error, and retire it if it doesn't.
Each stage is easy to state and hard to get right. Sections 1 to 4 cover the principles: why agents repeat mistakes, what makes a good error signal, how to read an agent trace, and why most failures turn out to be context failures. Sections 5 to 9 show a working implementation, the Entelligence Learning Loop, with numbers from an eight-week rollout. Sections 10 and 11 look at where the loop goes next.
The core idea is simple. Models get better every quarter, but no model will ever ship knowing your codebase's contracts, conventions and incidents. That knowledge has to be learned from your own factory's mistakes, and stored somewhere the next session can use it.
1. The problem: agents repeat themselves
Anyone who reviews agent-written PRs at volume sees this. Cluster a few weeks of substantive review comments and the same failures come back week after week, across different engineers, different repos and different agents.
The pattern is consistent. An agent passes user_id where the attribution contract expects cohort_id. It hand-rolls pagination offsets in a repo that has a paginate() helper. It calls db.session.execute_raw, a SQLAlchemy-shaped method that doesn't exist in the in-house ORM wrapper. A reviewer catches it, the engineer fixes it, the PR merges, and three days later a different session makes the identical mistake.
The reason is structural. Every agent session starts cold. A human engineer who gets that review comment once never makes the mistake again, because they remember. The agent has no equivalent of that memory. Static rules files (CLAUDE.md, AGENTS.md, .cursorrules) are the industry's current answer, but they are hand-written, rarely updated, never measured, and grow until agents stop reading them carefully.
In the rollout described later in this paper, 94% of substantive review comments across six repos traced back to a specific agent edit. These aren't human mistakes slipping through. They are agent mistakes, and they are learnable.
2. Code review is the error sensor
Every learning loop needs a reliable signal of failure. For coding agents, the best signal already exists: the code review. A review comment is a labeled example, written by an expert, pinned to a file and line, attached to a PR that an agent session produced.
To feed a learning loop, a review signal needs three properties that ordinary code review doesn't guarantee:
- Classified. Substantive issues (a broken contract, a wrong API, a violated convention, security, logic) are separated from nits. Nits stay out of the loop by default, because training agents against style preferences produces noise, not quality.
- Linked to the edit that caused it. A comment on
views.py L140has to join to the exact edit event in the agent trace (say, edit #612) that wrote that line. Without that link you know what's wrong, but not where the agent went wrong. - Inclusive of human reviewers. Comments from engineers, not just automated reviewers, flow into the same pipeline, so the loop learns the team's real standards.
Precision matters more than recall. A loop fed false positives teaches agents to avoid problems that don't exist, and spends context doing it. A factory's memory can only be as good as its failure signal.
3. Trace forensics: down to the step and turn
A review comment tells you what went wrong. The agent trace tells you how and why. The first step is to normalize sessions from every harness in use (Claude Code, Codex, Cursor and others) into a single event model.
Every event in a session is typed and indexed:
| Event type | What we capture | Why it matters |
|---|---|---|
| Prompt | User instructions and follow-ups | Did the human give the agent the constraint? |
| Exploration | File reads, greps, directory listings | What did the agent actually look at before editing? |
| Edit | File, line range, diff, event index | The line a reviewer later flags |
| Subagent | Spawn, calls, cost, handoff summary | What was learned in a subagent and lost on return |
| MCP / tool | Tool calls, arguments, results | Which context sources were queried, and what came back |
| Error | Failed tool calls, permission denials, test failures | Where the agent hit friction and improvised |
| Skill | Skills loaded and when | Whether available guidance was ever used |
The example session in this paper is prod-db agent PR review investigation plan, run on Claude Code for 4h 45m, 921 events, 8 subagents, 11 errors, $13.11. It produced PR #4182, which drew three substantive review comments and one nit.
For each comment we compute a causal path: the minimal chain of events that led from what the agent knew to the flagged edit. For the attribution bug, that path is four events long:
- #388, Sub 4 (Explore): reads
model_attribution.py, which defines the contract. - #402, Sub 4: returns a handoff summary to the main agent that omits
cohort_id. - #598, Main: reads
views.py. - #612, Main: edits
views.py L140, passinguser_id. Flagged in review.
The subagent had the right answer. The main agent never received it. You cannot see that from the diff; you can only see it from the trace.
We also compute ghost nodes: files that the fix required but the agent never opened. In this case attribution_serializer.py defines the contract and was never read. Ghosts are often the single most actionable output of the analysis, because they name exactly what context was missing.
4. Root cause: most failures are context failures
The model usually isn't the problem. When we classify causal paths, the large majority of substantive errors trace to the agent not having, not finding, or not keeping the right context at the moment it edited. Only a minority are genuine reasoning failures.
We classify every causal path into one of six root causes:
| Root cause | What happened in the trace | Example | Comments / session before learnings |
|---|---|---|---|
| Missing context | The fact the agent needed wasn't in anything it read | FE charts redeclare cohort types that exist in @/types/cohort | 0.38 |
| Ignored convention | The convention exists in the repo, but the agent never opened it | Read specs/, never CONVENTIONS.md; hand-rolled pagination | 0.31 |
| Lost subagent handoff | A subagent found the answer; the summary back to main dropped it | cohort_id contract read at #388, lost at #402 | 0.24 |
| Hallucinated API | The agent used a plausible method from public training data | execute_raw exists in SQLAlchemy, not in the house wrapper | 0.19 |
| Plan drift | The plan was right; execution diverged from it over a long session | Edits late in a 900-event session contradict the plan file | 0.17 |
| Tool / permission | A tool failed or was denied and the agent worked around it badly | Permission denied, then an Edit bypass that skipped a hook | 0.13 |
Three of these (missing context, ignored convention, hallucinated API) are fixed by supplying a fact. One (lost handoff) is fixed by changing how the harness moves context between agents. Plan drift and tool failures need timing: the right reminder, at the right step.
That split is the core design insight behind the memory layer. "Add more to CLAUDE.md" solves only the first category, and poorly. Each root cause needs a different delivery mechanism, which is why Section 7 treats where, when and how context is added as an experimental variable, not a constant.
From principles to practice: the Entelligence Learning Loop
The four stages above are a design pattern; any team can build toward them. At Entelligence we've built the full loop as a product, and the rest of this paper shows how it works and what it measured.
We started from the error sensor. Entelligence's code reviewer ranked #1 of 8 tools on real production bugs (47.2% F1) in our Code Review Benchmark 2026, which gave us a high-precision failure signal to learn from. From there we added trace ingestion, root-cause analysis, a shared memory we call the World Model, and the Entelligence CLI, which builds each agent session with the right learnings installed. As far as we know, no one else closes this loop end to end.
The Learning Loop has one view per stage: Pinpoint (one session), Aggregate (all sessions), Inject (the audit), and Prove (over time). Sections 5 to 9 walk through each. The numbers come from an eight-week rollout, Aug 4 to Sep 29, across 1,284 Claude Code sessions in six repos. The headline result: sessions built with learnings drew 53% fewer substantive review comments than a holdout, at $2.41 of injected context in total.
5. Where in the session each error was made, and why
For every flagged line, we show the engineer the exact moment in the session the error entered, the path that led there, and what the agent should have read. The session replay renders each review comment on the session timeline at the edit that wrote it, then draws the causal path back through subagents and reads.

The diagram below pulls out that one causal path so it can be read at a glance.
The subagent did its job; the handoff didn't. That distinction changes the fix: this isn't a rule to add to CLAUDE.md, it's a learning about how Explore subagents must report back (L-9), plus a path-specific fact about the attribution contract (L-14).
In the product, the replay also shows the session's full context tree: every file, MCP call, subagent and insight source the agent touched, sized by how often, with ghost nodes drawn dashed. Across the whole PR #4182 session, that tree covered 328 exploration events, 113 tool calls, 11 MCP calls and 11 errors.
6. From one session to recurring patterns
One causal path is an anecdote. The same causal path in five sessions is a pattern worth teaching. We cluster root-caused comments across every session in the org, by root cause, trigger path and the contract or convention involved, and propose a learning when a cluster crosses an evidence threshold.

Each learning carries its evidence trail: every PR comment it came from, with a link and a replay of the session that produced it. That trail is what a human reviews at the approval gate, and what makes each learning auditable later.
Across 1,284 sessions and 412 substantive comments over eight weeks, the loop produced six learnings:
| Learning | Rule | Root cause | Evidence | Delivery | Status |
|---|---|---|---|---|---|
| L-14 | Attribution in views goes through AttributionSerializer; expects cohort_id | Lost handoff + missing context | 9 comments · 5 sessions · 2 repos | World Model MCP | Active |
| L-11 | ORM wrapper has no execute_raw; use db.run_sql(query, params) | Hallucinated API | 5 comments · 5 sessions · 2 repos | Pre-edit hook | Active |
| L-12 | FE charts import cohort types from @/types/cohort | Missing context | 8 comments · 6 sessions · 1 repo | Rules file | Active |
| L-9 | Explore subagents return every file path and contract they read | Lost handoff | 7 comments · 6 sessions · 3 repos | Skill | Approved |
| L-15 | List endpoints use core.pagination.paginate() | Ignored convention | 6 comments · 4 sessions · 1 repo | Rules file | Proposed |
| L-6 | Name DataFrame variables with a _df suffix | Style preference | 11 nits · 3 sessions · 1 repo | Rules file | Retired |
Note L-9: it spans three repos because it isn't about any codebase. It's about the harness itself. Some of the most valuable learnings fix how agents work, not what they know.
7. Closing the loop: injection through the Entelligence CLI
A learning is worthless if it doesn't reach the next session at the moment it matters. The Entelligence CLI builds each agent session: before the harness starts, it resolves which learnings apply to this repo, this task and these paths, and installs them through one of four mechanisms.
| Mechanism | How it's delivered | Best for | Cost profile |
|---|---|---|---|
| Rules file | Compiled into CLAUDE.md / AGENTS.md at build | Short, always-true conventions (L-12: import cohort types from @/types/cohort) | Always in context; cheap per item, expensive in aggregate |
| Skill | Bundled as a skill, loaded when the task type matches | Procedural fixes (L-9: Explore subagents must return every contract they read) | Zero until the task matches |
| World Model MCP | Registered at build; retrieved when trigger paths match | Rich, path-specific knowledge (L-14: attribution goes through AttributionSerializer) | Pay only on retrieval, ~1.1K tokens |
| Pre-edit hook | Runs deterministically before edits to trigger paths | Hard constraints that must never be violated (L-11: no execute_raw, use db.run_sql) | ~0.3K tokens; enforced, not suggested |
Every learning carries a trigger, a path glob plus a diff condition (for example backend/**/views.py · diff touches attribution). Triggers are what keep context lean: across 1,027 sessions built with learnings, the CLI made 2,356 insertions at an average of just +0.9K tokens per session.

What we experiment on
We treat injection as a four-variable experiment, and run it continuously:
- What to add: the rule itself, the evidence behind it, a pointer to the file that defines it, or all three. Pointers to ghost files often outperform prose rules.
- Where to add it: rules file, skill, MCP or hook. The same learning can succeed as a hook and fail as a rules line.
- When it fires: at session build, at task-type match, at first read of a trigger path, or immediately before an edit. Later is usually better; the agent acts on what's in its recent context.
- How much: per-session token budget, ranking when many learnings match, and suppression of learnings that have stopped paying off.
The audit log keeps us honest about delivery versus use. A learning can be delivered and still ignored. In PR #4214, L-14 was retrieved, then ignored, and the error recurred. In PR #4182, the L-11 hook never fired because the agent wrote execute_raw through an Edit path that bypassed it. Both are logged as failures of the mechanism, not successes of delivery, and both feed back into the next round of experiments.
Governance
Learnings move through an explicit lifecycle: Proposed (clustered from review comments) → Approved (human gate; nothing enters a session without approval) → Active (added by the CLI) → Retired (removed when it stops paying off). Engineers stay in control of what their agents are told.
8. The World Model: where the memory lives
The World Model is a graph of everything the loop has learned about how agents behave in your org. It stores every session, PR, review comment and person as nodes, and the causal links between them as edges. In the view below, the top 238 patterns, people and themes are drawn from a graph of 162,130 nodes.

In this loaded view, the graph holds:
| Node type | Count in view | What it represents |
|---|---|---|
| Issue | 429 | A specific failure, e.g. "Stale API Key Retried on 401", "Localhost Fallback in Prod" |
| Theme | 150 | A cluster of related work or failures, e.g. "Ellie CLI audit and merge" |
| Session | 132 | An agent trace, with cost, e.g. "claude-code · $376" |
| Pattern | 69 | A recurring root cause across sessions, e.g. a lost subagent handoff |
| PR | 40 | The diff a session produced, e.g. PR #5220 |
| Person | 19 | The engineers whose sessions feed the graph |
The view has three lenses: Recurring Issues (what keeps breaking), Coding Agent (how agents behave), and Merged (both, linked). 283 links connect the loaded nodes, and hovering an orb traces them. Ask Ellie answers questions against the same graph in natural language.
The schema is what makes the memory trustworthy. Every learning resolves back through a pattern to specific issues, to the edits that caused them, to the sessions and people involved. Nothing an agent is told is unsourced.
Agents read the graph live through the world_model.search MCP tool, which the CLI registers at build. That is the delivery path for rich, path-specific learnings like L-14: rather than bloating every session's rules file, the agent retrieves the relevant slice of memory only when it touches backend/**/views.py.
9. Proving it: the same errors, measured over time
Sessions built with learnings drew 53% fewer substantive review comments than the holdout (0.62 vs 1.31 per session; 95% CI −41% to −63%). Every claim of improvement in this system is made against a control, not against last month.

How we measure. Since Aug 4, 20% of sessions are a standing holdout. For those sessions, the CLI still matches triggers and logs which learnings would have been added, but adds nothing. That gives us a like-for-like control: same repos, same engineers, same agents, same weeks. The holdout itself improved about 10% over the period, which is the team-effect baseline (people get better, models get better). We report the gap above that line, not the raw drop.
The lines separate the week learnings switch on and keep separating as more learnings go live. The holdout barely moves, which is the point: the improvement comes from the memory, not from the model or the calendar.
Because every comment is root-caused, we can also see which kinds of failure the memory fixes.
This is the honest part of the result, and the one we find most useful. Context injection crushes failures caused by missing facts and lost handoffs. It does much less for plan drift, which is a long-horizon execution problem. That tells us where to push next: mid-session reminders timed to plan checkpoints, rather than more context at build.
| Metric | With learnings | Holdout |
|---|---|---|
| Sessions | 1,027 | 257 |
| Review rounds to merge | 1.3 | 1.9 |
| Agent errors per session | 4.1 | 6.8 |
Per learning, not just in aggregate. Each learning gets its own scorecard, so a weak learning can't hide behind a strong one:
| Learning | Delivery | Triggered | Used | Recurred | Prevented | Cost / session | Verdict |
|---|---|---|---|---|---|---|---|
| L-11 | Pre-edit hook | 64 | 58 | 3 | 14 | $0.001 | Keep |
| L-12 | Rules file | 41 | 33 | 4 | 8 | $0.001 | Keep |
| L-14 | World Model MCP | 38 | 35 | 2 | 9 | $0.002 | Keep |
| L-6 | Rules file | 44 | 26 | 18 | 1 | $0.001 | Retired |
L-6 is the loop working as designed. It was a style rule learned from nits; it recurred 18 of 44 times, prevented almost nothing, and was retired. Memory that doesn't pay for itself gets removed.
Cost. Injection averages +0.9K tokens and $0.0023 per session: $2.41 across eight weeks. L-14 alone costs $0.002 per session and prevents about one review round-trip every eight sessions. That's the economic case in one line: a fraction of a cent of context replaces an engineer's review cycle.
10. Beyond code review: every signal your org already has
Code review is the first sensor, not the only one. The same loop works for any signal that says "an agent-written change caused a problem", and on the upgraded Entelligence plan these plug into the World Model today:
| Source | What it teaches the agent | Example learning |
|---|---|---|
| Production incidents (PagerDuty, incident.io) | Which changes broke prod, and the fix that resolved them | "Migrations on large tables must be batched; an unbatched one locked the table in prod" |
| Alerts and errors (Sentry, Datadog) | Runtime failures traced back to the PR and session that shipped them | "This endpoint's null team_model crashes in prod; guard it" |
| Designs and specs (Figma, Notion, Linear) | Intended behavior and decisions the code doesn't show | "Checkout copy and states follow the approved Figma flow" |
| Customer tickets (Zendesk, Intercom, Slack) | Real user pain linked to the code that caused it | "Large exports time out for enterprise customers; stream them" |
Every source is normalized the same way: an event, linked to code, linked to the session that wrote it, clustered into a learning, gated by a human, delivered by the CLI, and measured by holdout. The loop is source-agnostic by design.
This is the long-term shape of the product: a production-aware memory for every coding agent in the org, so that the incident your on-call engineer fixed at 3 a.m. is something no agent in your company will ever cause again.
11. Open source, and where this goes
The world model visualization is open source: Entelligence-AI/agentplay. Point it at your own agent sessions to see the patterns, people and themes in your graph, and how they connect.
We're open-sourcing the view because we want the whole industry looking at agent traces this way. Agents will keep getting smarter. What they won't get on their own is the memory of your codebase: the contracts, conventions, incidents and decisions that make your code yours. We are building the layer that captures that memory, delivers it, and proves it works. As far as we know, no one else closes this loop end to end, and we intend to define what it looks like.