6.0 Toward a Theory of MAGE
The book has argued from one case, seen in depth. A theory is what lets that case speak past itself: it names the constructs, draws the mechanisms between them, and states what evidence would make the account fail.
The theory offered here is middle-range. It does not claim that every software project will behave like the one this book studied, and it assigns no universal coefficients to the relations it draws — a single system cannot support them. What it offers is a dynamic model of a specific phenomenon: how a fleet of coding agents turns abundant implementation capacity into durable engineering progress, or into churn. Read the whole chapter with that scope in mind; I will not restate it under each claim.
The central claim is not that agentic velocity is inherently productive. Velocity amplifies the fit between the fleet and the engineering environment it runs in. When that environment is incomplete, incoherent, or too weak to hold the work, velocity produces churn, escaped defects, repeated intervention, and governance friction. Those signals create pressure to adapt the environment. Engineers and engineering organizations respond by modeling missing structure, converting recurring failures into durable mechanisms, and reconciling or retiring governance that has grown excessive. When the adaptation works, later agents inherit a more legible action space, and a greater share of raw velocity becomes durable progress.
6.0.1 The dynamic model
Figure 6.0-1 draws the whole theory on one diagram: a single set of constructs rather than four separate pictures, because the four dynamics act on the same system. The figure follows the causal-loop convention of system dynamics 11. Donella H. Meadows, Thinking in Systems: A Primer (Chelsea Green Publishing, 2008)..
The spine runs left to right. Agentic capacity — model capability, concurrency, raw implementation rate — exercises a governed environment whose effective quality the model calls E. That environment shapes agentic work into realized performance: durable throughput, defect escape, and the human attention each unit of durable output costs. The gap between the performance the fleet looks capable of and the durable performance it delivers registers as governance pressure. Pressure meets structural diagnosis — the judgment that reads a signal as a local defect or as evidence of missing structure. Diagnosis drives governance adaptation, which changes E. The changed environment reopens the loop.
Four numbered dynamics run over that spine, and each has a color and a number so the figure survives in grayscale:
- Blue — the bootstrap path. Prior engineering knowledge seeds the initial environment before any agent runs.
- Red — the churn/discovery loop. Work outruns the environment; churn and failure make the missing governance visible.
- Green — the governance-adaptation loop. A structural diagnosis reshapes the environment for the next agent.
- Purple — the capability-amplification loop. A stronger environment raises ambition and concurrency, which exposes a new frontier.
Follow the red loop first. It carries the experienced driver of the whole system — the pressure an engineer actually feels — and the other three are best read as answers to it: blue lowers the pressure before it starts, green relieves it once it appears, and purple explains why it never stays relieved.
The four dynamics are not all the same kind, and naming their types sharpens why the system does not settle. Three are feedback loops; bootstrap is the input that sets the initial condition. Loops [2] and [3] together form a balancing loop — churn drives diagnosis, diagnosis drives adaptation, and adaptation reduces churn: a goal-seeking cycle that drives the environment toward fit. Loop [4] is a reinforcing loop — a better environment raises ambition, and greater ambition reopens the frontier: a self-amplifying cycle that keeps moving the target. The balancing loop settles the environment toward fit; the reinforcing loop moves the target the moment it is reached. That pairing — balance toward fit, then reinforce to reopen the frontier — is why MAGE never converges.
6.0.2 The four dynamics
Each dynamic gets its own section below, and each opens with a simplified diagram of just its local causal path. The master figure above already established that the four are one system, so each section shows only the fragment it explains.
The blue path: bootstrap governance
Not all governance is discovered through failure; the blue path, in Figure 6.0-2, seeds the environment before any agent runs.
A project begins with knowledge it inherited and knowledge it can anticipate. A brownfield system already carries architecture, tests, deployment controls, operational history, and institutional obligations. A greenfield project starts with less local evidence, but a skilled engineer can still seed it from prior experience, external standards, known security and reliability practices, and lessons transferred from earlier systems.
Bootstrap governance converts that ex-ante knowledge into the initial governed environment: a known obligation or a prior lesson becomes a model, a mechanism, or a constrained architecture, and the sum of them sets the environment's starting quality. This is why disciplined organizations and experienced engineers need not rediscover every failure from scratch. Governance conversion in one project becomes bootstrap knowledge in the next.
Bootstrap governance is never complete. The environment a product needs depends on the product, its trajectory, and the agents acting on it, and none of that is fully knowable in advance. The blue path lowers the initial failure rate. It does not close the red loop.
The red loop: churn makes missing governance visible
The red loop, in Figure 6.0-3, is where work outruns a weak environment and the gap surfaces as pressure.
Agentic velocity is not the immediate motive for governance investment. The motive is the experienced gap between what the fleet appears capable of producing and what becomes durable progress. When the environment is weak, raw implementation capacity produces repeated rediscovery, regressions and reversals, escaped defects, recurring human intervention, control friction, and confidently wrong changes that consume tokens without advancing the system. The book calls the dominant form churn.
Velocity earns its place in the model as an amplifier. It clusters latent weakness in time: failure classes that might have arrived months apart under human-paced work land in a single afternoon, and their proximity makes the structure they share visible. Governance pressure is what the operator feels when that happens.
Pressure is not yet knowledge. The next act is judgment. Structural diagnosis distinguishes a local defect from a structural failure class. A local defect gets repaired and the matter ends. A structural failure reveals a weak or absent model, abstraction, boundary, invariant, oracle, mechanism, or governance seam — something whose absence will keep producing failures until it is supplied. This diagnosis is the scarce work of the whole method, and the red loop is where it is demanded.
The green loop: governance adaptation
A structural diagnosis triggers governance adaptation, the green loop of Figure 6.0-4.
The core move here is governance conversion: turn a recurring failure class into a durable mechanism the environment enforces on every later agent. Mature adaptation keeps that move and adds four more around it:
- Model — make missing intent or structure legible, so agents reason over the right representation.
- Convert — turn a recurring failure into a durable mechanism.
- Strengthen — replace weak probabilistic guidance with harder enforcement where the obligation justifies it.
- Reconcile — resolve controls that duplicate, conflict, or fire in the wrong order.
- Retire — remove governance that has gone stale or become net-harmful.
The output is not merely more controls. It is a change in the shape of future work. Later agents inherit a more compact representation, clearer constraints, fewer invalid moves, earlier evidence, narrower sanctioned seams, and more reliable admission conditions. That inheritance is the immediate mechanism by which the environment's quality changes — and it is why adaptation can be subtractive as well as additive: reconciling a control collision or retiring a stale gate raises quality without adding a thing.
The quality that adaptation moves needs a definition sharp enough to keep apparatus quantity from standing in for it.
Effective governed-environment quality, E. The degree to which the environment gives agents a current, task-relevant representation of the system, together with coherent, proportionate, machine-actionable coverage of its important obligations.
That quality has four parts, set out in Table 6.0-1; each asks a question the raw apparatus count cannot answer.
| Dimension | The question it asks |
|---|---|
| Representation quality | Are the relevant properties legible in the environment's current models? |
| Enforcement coverage | Are the important known obligations held mechanically, not merely watched? |
| Coherence | Do the mechanisms compose without harmful collision or duplication? |
| Economy | Is the governance burden proportionate to the risk and value it controls? |
Defined this way, E is a quality, not a quantity. A large support apparatus raises E only if it is coherent, economical, and actually covers the obligations that matter; the same apparatus can instead be duplicated machinery, stale enforcement, or a tower of governance that lowers E while the control count climbs. Lint count and support ratio are indicators of investment. They are not substitutes for quality.
The purple loop: capability amplification
Successful governance does not end the process; the purple loop of Figure 6.0-5 reopens it.
It expands what the engineer attempts. A stronger environment makes larger delegated tasks feasible, supports greater concurrency, invites more ambitious product requirements, cheapens refactoring, and lets the operator run agentic capacity denser. Every one of those pushes the work into surface the environment has not yet governed, and the expanded frontier exposes new weaknesses.
This is capability amplification, and it is why MAGE does not terminate. The environment never converges on a final set of controls, because the product, the fleet, and the task frontier keep changing underneath it — and each gain in governance raises the ambition that finds the next gap. The loop also answers a natural objection: that better agents will eventually remove the need for governance. They do the opposite. A more capable fleet raises attainable velocity and ambition, which raises the consequences of a weak model or a missing control. The frontier moves out with the capability, and governance follows it.
6.0.3 Constructs and path-specific moderators
The two conditions the book leaned on — capability-fit and authority — are not global switches over the whole model. Each acts on one specific transition, and stating where sharpens what each one predicts. Table 6.0-2 places every construct and moderator on the transition it governs.
| Construct or moderator | Role in the model | Operates on |
|---|---|---|
| Agentic capacity | Model capability, concurrency, raw implementation rate | Input to agentic work |
| Governed-environment quality E | Effective quality of models and mechanisms, not their quantity | Work → realized performance |
| Governance pressure | Churn, escapes, intervention burden, control friction | Performance gap → diagnosis |
| Structural diagnosis | Reading a failure as local, or as evidence of missing structure | Pressure → adaptation |
| Governance-adaptation capability | Capacity to model, convert, strengthen, reconcile, and retire | Diagnosis → environment change |
| Capability-fit | Fit between the work, agent capability, and the judgment a diagnosis needs | Moderates pressure → diagnosis |
| Authority | Ability or organizational path to modify shared architecture and controls | Moderates diagnosis → adaptation |
| Domain governability | Degree to which important properties can be represented and mechanically checked | Moderates E → outcomes |
| Control-estate coherence | Degree to which mechanisms compose without harmful interaction | Moderates apparatus → E |
Capability-fit sits on the pressure-to-diagnosis step. A capable fleet given to someone who cannot tell a local defect from a structural one produces motion without direction; the pressure arrives, but no diagnosis converts it. Authority sits one step later, on diagnosis-to-adaptation. The person who reads the recurring failure has to be able to change the environment — the architecture, the mechanisms, the gates — or the diagnosis dies as a local patch and the class never dies with it.
Authority need not live in a single architect. In an organization it is a functioning path from an observed failure to shared infrastructure. Split ownership across teams, owners, and review boards is not fatal on its own. The failure comes when the path terminates in a local patch: the same signal that could have become a shared mechanism produces a private fix instead, and the class survives to recur under the next owner.
6.0.4 A capability model, not a maturity model
MAGE is a capability model, not a maturity ladder, and the distinction matters because the field has been here before. For two decades the dominant frame for engineering quality was the staged maturity model — the CMM and its CMMI successor — which ranked an organization on a five-rung ladder and told it to climb. A staged ladder imposes one linear progression on every organization regardless of where its actual constraint sits, grades process conformance in place of the outcomes the process was meant to produce, and under pressure decays into box-ticking, where reaching a level stands in for shipping better software 22. Terry B. Bollinger and Clement L. McGowan, “A Critical Look at Software Capability Evaluations,” IEEE Software 8, no. 4 (1991): 25–41, https://doi.org/10.1109/52.300034.. The objection is not to structure. It is that a ladder makes the rung the goal.
The dynamic model gives the alternative directly. A project's current bottleneck selects the next modeling or governance capability it needs; the bottleneck in front of you picks the next control, and no rung waits above it. When escapes climb, you reach for the artifact-side controls. When agents churn against a surface they cannot hold, you invest in the models. When recovery keeps pulling a human back in, you harden the fleet's substrate. Dependencies may impose a local adoption order — some controls presume others — but that order is a dependency, not a rank, and no universal sequence ranks every organization. The reason there are no levels is the reason there are no fitted coefficients: a single case cannot validate a progression that means the same thing everywhere. Different projects enter with different bootstrap environments and feel different pressures, so what replaces "what level am I" is a plainer question — which capabilities do you have, and what does each one buy against the bottleneck you actually face.
6.0.5 Outcomes and observables
The model earns its keep only if its outcomes can be watched. Three of them carry the theory, and they are the three dimensions the productivity literature already names — velocity, quality, and sustainability — read for a fleet rather than a team. Table 6.0-3 states each and the observables that would track it.
| Outcome | What it measures | Example observables |
|---|---|---|
| Durable throughput | Raw implementation converted into lasting system progress | landed changes net of rework; churn share; the slope of sustained throughput |
| Defect escape | Relevant failures reaching later lifecycle stages or production | escaped-defect rate; reopen rate; silent-corruption escape |
| Human-attention burden | Human judgment required per unit of durable output | interventions per landed change; review time; conversion effort |
State human-attention burden as a rate, not a ceiling. The theory predicts amortization: as recurring failure classes convert into durable governance, the attention required per landed change falls, even if total governance work continues to grow. Per-change attention declines; the aggregate need not.
Apparatus measures are indicators, not quality
A second discipline governs how the environment itself is read. The quantities easiest to count — the support ratio, the model-coverage fraction, the control count, the gate count, the model-sync rate, the control-interaction findings, the stale-control rate — measure investment or accumulation. They do not measure whether the environment is any good.
A large support apparatus can be effective governance. It can also be duplicated machinery, accidental complexity, stale enforcement, a tower of governance, or an internally conflicting control estate. The number cannot tell you which. A high support ratio gains meaning only when it is joined to an outcome — when the apparatus that leads production is shown to sustain throughput or drive escape down. Treat every apparatus measure as an input or an indicator. Effective governed-environment quality is what the outcomes reveal, and the apparatus is only its raw material.
6.0.6 Principal predictions
Here is where the model sticks its neck out. The dynamic model yields seven principal hypotheses. Each names a mechanism and the single observation that would sink it. Core causal propositions appear here; derivative consequences and domain-specific instances follow as corollaries and specializations.
None of the seven is tested by this case. One repository cannot test a claim about a population. Every row is offered so a larger study can, and every row is stated as a directional, conditional claim — the shape it should take, and the conditions under which it should hold — not a measured law. Table 6.0-4 lists the seven with the observation that would falsify each.
| ID | Hypothesis | Key falsifier |
|---|---|---|
| H1 — Failure-class exposure | Higher agentic velocity shortens the elapsed time to recognize recurring failures as one shared structural class, controlling for the opportunities to trigger the class. | Recognition latency does not improve with velocity once exposure opportunities are controlled. |
| H2 — Governance moderation | Governed-environment quality sets the sign of velocity's effect: higher velocity raises churn in low-quality environments and durable throughput in high-quality ones. | Velocity relates to durable throughput the same way regardless of environment quality. |
| H3 — Mechanized assurance | Greater risk-weighted coverage of relevant failure classes by effective deterministic mechanisms lowers defect escape, and its advantage over attention-based control widens as velocity rises — for adequately scoped mechanisms. | Mechanized coverage adds no protection, or its advantage over review does not widen as human attention saturates. |
| H4 — Representation leverage | H4a — Representation efficiency. Task-relevant structured models reduce context cost and churn without lowering task quality. H4b — Synchronization signal. Model-synchronization measures provide predictive information about later churn or defect escape beyond conventional code-coverage measures. | H4a: Model use at matched recall yields no efficiency gain or lowers task quality. H4b: Synchronization measures add no predictive power over conventional code-coverage measures. |
| H5 — Local model deficiency | Unmodelled or drifted architectural surface predicts later churn or escape in the same subsystem. | Orphaned or drifted surfaces are no more likely than modelled ones to see later churn or escape. |
| H6 — Oversight amortization | As recurring failure classes convert into durable governance, human attention per unit of durable output declines — even if total governance work keeps growing. | Interventions per durable change stay flat or rise as effective conversion coverage grows. |
| H7 — Conversion conditions | Structural failures convert into shared governance faster and more durably when the capability to diagnose structure and the authority to change the environment are colocated, or connected by a working organizational pathway. | Split-authority or low-fit settings convert at the same rate and durability as high-fit, effective-authority ones. |
Two of the seven carry conditions worth stating in the open. H3 does not claim hard controls never fail; it claims that a correctly encoded, adequately scoped mechanism is less sensitive to throughput than finite review, so the gap between them opens as the fleet speeds up. H6 does not promise the total oversight bill stops growing; it claims the per-change share falls as classes convert. Strip those conditions and both claims overreach; stated with them, each says only what the evidence can bear.
Corollaries and specializations
Several claims read more precisely as consequences or domain-specific instances of the seven than as coequal predictions. Naming them as such is not a demotion of their truth. It is a statement of where each sits in the causal structure; Table 6.0-5 places them.
| Former prediction | New status |
|---|---|
| Speed and stability do not trade | Corollary of H2 and H3 |
| Conversion gets faster as tooling matures | Learning-curve specialization of H7 |
| The substrate heals faster as it accretes controls | Consequence of H3 and H6 |
| Grammar coverage beats line coverage | Domain-specific instance of H4 |
| A fidelity gate reduces silent corruption | Product-specific instance of H3 |
| Deployment pain stays bounded | Subjective consequence of H6; future work |
The speed–stability claim, the one Accelerate made famous, now falls out of H2 and H3 together rather than standing alone: in a high-quality environment, pushing frequency up does not push the change-fail rate up. The two domain instances — grammar coverage driving out format escapes, a content-fidelity gate killing the silent-corruption class — are H4 and H3 read on one product's surface, valuable as grounding but not as separate laws.
6.0.7 Two directional relations
Two of the hypotheses have a form clean enough to write down. Neither is a fitted law. Each is a shape — a statement of which way a quantity moves and why, with the parameters named but not measured. Read them as the arithmetic behind the arrows, not as equations with coefficients this case supplied.
The churn relation
Durable throughput is raw velocity net of what churn eats:
V_durable = V_raw · (1 − c)
where V_raw is the raw agentic implementation rate and c is the share of that effort consumed by churn. The substantive claim is about what moves c. As effective modeling coverage rises, more of the relevant surface fits the fleet's window before it churns, so over the under-governed range
∂c / ∂E < 0
churn share falls as environment quality climbs. That is the arithmetic behind H2 and its subsystem sharpening in H5.
But the relation is not monotone across all governance accumulation, and the honesty of the theory turns on saying so. Push apparatus past the point of fit and churn starts to rise again — from collisions, false positives, latency, context burden, and constraints that block legitimate work. The qualitative shape has three bands, drawn in Figure 6.0-6: churn high in the under-governed region, lowest at fit, rising once more where the estate overgrows into a tower of governance.
under-governed → churn high · well-fit → churn lowest · overgrown → churn rises from friction
No exact curve is claimed. The point is the non-monotonicity: apparatus quantity is not the same thing as effective environment quality, and a theory that assumed more governance is always better would miss the entire right-hand band.
Attention-based versus mechanized control
The second relation contrasts two ways to hold a failure class. Let p be the risk-weighted rate of defect opportunity, r(V) the fraction a finite human review still catches at velocity V, and κ the risk-weighted fraction of relevant opportunities covered by effective mechanisms. Then the two escape rates are:
ε_review ≈ p · (1 − r(V)) and ε_control ≈ p · (1 − κ)
Human attention is finite, so r(V) falls as velocity rises — review catches a smaller share of a larger stream, and ε_review climbs toward the raw defect rate p. A mechanism applies to every change regardless of throughput, so κ does not fall with velocity in the same way. The claim is not that hard controls never fail. It is conditional and comparative: for adequately scoped mechanisms, κ is less sensitive to throughput than review coverage r(V), so the daylight between "looked obeyed" and "was enforced" widens as the fleet speeds up. That is H3 in one line.
6.0.8 Evolution of this theory
The theory did not arrive whole. Its empirical core came first, induced from a bounded case; the rest is later reflection, and honesty requires keeping the two apart.
The first version appeared in the case-study paper Cheap Code, Costly Judgment 33. James C. Davis et al., “Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering,” 2026, https://arxiv.org/abs/2607.01087.. That paper analyzed the first twelve weeks of the DocAble build and induced a five-step governance-conversion loop:
- velocity exposes failure;
- the engineer classifies the failure as local or structural;
- structural failures are converted into governance;
- later agents inherit a narrower, more explicit action space;
- repeated conversions raise the environment's capacity to absorb future work.
That loop remains the empirical kernel. It contributes distinctions this chapter preserves: failure must be interpreted, not merely observed; governance works by reshaping subsequent action; conversion is recursive and does not terminate; capability-fit and authority condition whether a structural signal becomes shared infrastructure.
Continued development and roughly two further months of reflection extended the kernel in four ways.
- Ex-ante governance became a first-class path. The paper emphasized the delta of ex-post governance — controls discovered from failures no one could fully anticipate. Later work sharpened the complement: a disciplined organization or skilled engineer bootstraps a project with inherited architecture, standards, tests, prior incidents, and lessons carried from earlier systems. Conversion in one project becomes ex-ante knowledge in the next. Bootstrap and conversion are two loops, not one.
- Modeling became a coequal route to governability. The paper treated structured models as one governance mechanism among others. The Modeling Thesis gave them a distinct role. Models do not only constrain output; they compress intent and system structure into a representation the fleet reasons through, surface properties at the right semantic level, and support the drift gate that holds the map equal to the territory.
- Governance accumulation proved non-monotonic. The paper saw a large, layered substrate and described repeated conversion as compounding governability, already noting the tendency was not strictly monotone. Continued work supplied the missing negative side: mechanisms collide, duplicate, consume context, and block legitimate change. Mature adaptation therefore includes reconciliation and retirement, not only accumulation.
- The process became visibly multi-loop. The early theory centered one loop — failure, judgment, governance. The book revealed interacting loops: bootstrap establishes the initial environment, churn reveals what is missing, adaptation reshapes future work, and success expands ambition and concurrency onto a new frontier.
The result is still a governance-conversion theory at its core. It now describes MAGE as an adaptive system for governing the engineering environment itself.
Epistemic note
The induced kernel and its later extensions are epistemically different, and the chapter should not blur them. The five-step loop was induced from the bounded case and its incident corpus. The four extensions are grounded in continued development, further observation, and integration across the book. Read the kernel as the case's direct yield and the extensions as refinements proposed for later testing — not as independently confirmed, population-level laws.
6.0.9 Research agenda
The seven hypotheses suggest a focused program, not a scatter of disconnected tests. Three study designs would cover most of the model, and none needs this repository; Table 6.0-6 maps each design onto the hypotheses it would test.
| Study design | Principal hypotheses |
|---|---|
| Multi-project longitudinal panel with environment-quality variation | H2, H3, H4, H6 |
| Subsystem-level panel tracking model gaps, churn, and escapes | H1, H5 |
| Organizational comparison of diagnosis and authority pathways | H7 |
A particularly strong design would instrument a failure class from first occurrence through its whole life — recurrence, structural recognition, governance intervention, and later recurrence or escape. Such an instrument would test the discovery and adaptation loops directly, rather than leaning on aggregate repository metrics that cannot separate "the environment matured" from "the operator did."
6.0.10 Closing
The claim of MAGE is not that more governance produces better software. It is that abundant implementation capacity makes the fit of the engineering environment decisive. Known obligations can be encoded before the work begins. Unknown ones become visible only through churn, failure, and friction. Judgment separates the local defect from the missing structure; authority lets that diagnosis alter the shared environment; models and mechanisms reshape what the next agent can know and do. When the loop works, velocity buys durable progress. When it does not, velocity only reaches the churn wall sooner. That is the theory this book leaves for the next case to test.
Works Cited
- Meadows, Donella H. Thinking in Systems: A Primer. Chelsea Green Publishing, 2008.
- Bollinger, Terry B., and Clement L. McGowan. “A Critical Look at Software Capability Evaluations.” IEEE Software 8, no. 4 (1991): 25–41. https://doi.org/10.1109/52.300034.
- Davis, James C., Paschal C. Amusuo, Tanmay Singla, Berk Çakar, and Kirsten A. Davis. “Cheap Code, Costly Judgment: A Case Study on Governable Agentic Software Engineering.” 2026. https://arxiv.org/abs/2607.01087.