6.1 A Theory of MAGE
Part V ended with two complementary views of agentic engineering: one system observed deeply enough to reconstruct mechanisms and sequence, and independent systems observed broadly enough to expose variation. This chapter proposes an explanation for both.
The theory is directional rather than numerical. It identifies constructs, mechanisms, and relationships that should hold under stated conditions. Section 6.3 turns those claims into testable predictions.
The central claim is simple. Agentic capacity acts through the engineered environment that surrounds it. A strong environment turns more implementation capacity into durable progress; a weak or incoherent one turns the same capacity into churn, escaped defects, repeated intervention, and control friction. Those failures expose possible gaps; engineering judgment diagnoses them; adaptation changes what later work inherits—the representations, mechanisms, or architecture. As the environment improves, larger tasks and greater autonomy become feasible, exposing the next frontier.
Theory of MAGE — informal
Agentic capacity does not by itself produce engineering progress. It magnifies the fit—or misfit—between the work attempted and the engineered environment that surrounds it.
6.1.1 The Dynamic Model
Figure 6.1-1 puts the theory on one diagram. Agentic capacity supplies model capability, concurrency, and raw implementation rate. The governed engineering environment supplies the representations, constraints, evidence, and authority through which that capacity acts. Together they shape realized performance: durable throughput, defect escape, and human attention consumed per unit of useful output.
A gap between apparent capacity and durable performance registers as governance pressure—churn, recurring failure, escaped defects, repeated intervention, or friction introduced by the governance itself. Pressure is evidence, not diagnosis. An engineer or organization still has to decide whether the signal is a local defect, a missing representation, a weak boundary, an inadequate control, or an obsolete piece of machinery. That diagnosis drives environmental adaptation, which shapes the next round of work.
Four dynamics govern this process. Bootstrap carries prior engineering knowledge into the environment before work begins. Discovery exposes mismatches between the work attempted and the environment available to support it. Adaptation changes the environment in response. Capability amplification follows successful adaptation: stronger environments support larger tasks, greater concurrency, or more ambitious requirements, pushing work onto a new frontier. The loop continues while capability, ambition, or the system itself changes.
The feedback structure also has a cybernetic heritage. Cybernetics studied how regulators use information about a system to keep consequential behavior within acceptable bounds. Conant and Ashby's good regulator theorem connects regulation directly to representation: under its assumptions, the simplest optimal regulator must embody a model of the system it regulates 11. Roger C. Conant and W. Ross Ashby, “Every Good Regulator of a System Must Be a Model of That System,” International Journal of Systems Science 1, no. 2 (1970): 89–97, https://doi.org/10.1080/00207727008920220.. MAGE adopts a related engineering intuition for autonomous software engineering. A governed engineering environment represents consequential properties of the system and its engineering process, observes what happens as work proceeds, and adapts when the available representations and controls prove inadequate. The regulated object differs, but the same engineering question recurs: what representation is adequate for the behavior being governed?
Not all useful structure is discovered through failure. Existing architecture, standards, requirements, operating history, and lessons imported from earlier systems can seed the environment before an agent acts. Other needs become visible only in use. MAGE therefore has two routes to durable structure: ex ante, where a known obligation or useful representation is engineered deliberately, and ex post, where experience exposes a recurring cost or risk worth externalizing. Governance conversion names the latter route; it is not the origin story of every model or mechanism.
Two conditions determine whether adaptation reaches the shared environment. First, someone has to recognize that pressure is structural rather than local. Second, that diagnosis needs a path to the shared environment—the architecture, models, mechanisms, or infrastructure that can retire the failure class. Without diagnosis, pressure produces motion without direction. Without authority, the lesson dies as a local patch. Section 6.2 treats both as scope conditions.
6.1.2 What Makes an Environment Effective
Apparatus quantity is not a proxy for environment quality. Ten thousand lines of stale checks are not better than one hundred lines that cover the obligations that matter. Call the relevant construct environment quality: the degree to which the governed engineering environment provides useful, current representations and proportionate coverage of the obligations relevant to the work.
Four dimensions make the definition concrete. Representation: are the properties needed for the work explicit at an appropriate abstraction? Coverage: are important obligations supplied with adequate evidence and authority? Coherence: do the representations and mechanisms compose without contradiction or harmful duplication? Economy: does the future value they produce justify their friction, maintenance, and context cost? Together they define environment quality; they do not reduce to a scalar score.
Coverage does not require mechanizing every obligation. Some obligations belong in a compiler rule or admission gate; others are statistical, semantic, or consequential enough that human judgment remains the appropriate authority. A strong environment does not mechanize everything. It makes explicit where judgment belongs and where repeatedly paying for it would be wasteful.
MAGE is a repertoire of capabilities rather than a maturity ladder. Investment follows the bottleneck: strengthen representation where reconstruction dominates; strengthen evidence or authority where failures escape; retire machinery where governance itself becomes the tax. Local dependencies may constrain the order of adoption without defining a global maturity level.
6.1.3 Outcomes and Observables
A theory of agentic engineering should distinguish activity from progress. MAGE therefore treats raw implementation velocity as an input and distinguishes three outcomes: durable throughput, defect escape, and human-attention burden. Table 6.1-1 defines each outcome and identifies candidate observables.
| Outcome | What it measures | Candidate observables |
|---|---|---|
| Durable throughput | Raw implementation converted into lasting system progress | landed changes net of rework; churn share; the slope of sustained throughput |
| Defect escape | Relevant failures reaching later lifecycle stages or production | escaped-defect rate; reopen rate; silent-corruption escape |
| Human-attention burden | Human judgment required per unit of durable output | interventions per landed change; review time; conversion effort |
Durable throughput is useful system progress that survives rework, integration, and operation rather than merely landing in the repository. Defect escape records consequential failures that reach later lifecycle stages or production. Human-attention burden records the scarce engineering judgment consumed per unit of durable progress.
Several easy-to-count quantities measure activity or investment rather than outcome: commits, lines changed, control files, model-sync events, and support ratio. They can help explain what an environment is doing; they do not establish that it is producing a return.** The activity-versus-outcome gap already shows up in industry data. A 2026 benchmark of 83,000 developers across 253 organizations 22. LinearB, “The Engineering Productivity Gap: How Elite AI Teams Are Pulling Away from the Rest,” LinearB, 2026, https://linearb.io/resources/ai-engineering-productivity-gap. reports large merge-rate gains among heavy AI users, yet finds yield falling among the heaviest even as merge rates roughly doubled; an industry analyst reads the same period as rising AI spend without proportionate shipping velocity, the difficulty having moved from comprehension to confidence 33. Jennifer Riggins, “AI Coding Got Faster. Why Didn't Engineering?,” The New Stack, August 9, 2026, https://thenewstack.io/ai-productivity-measurement-gap/.. Activity is not the same as durable progress.
6.1.4 Three Core Predictions
Three predictions form the core of the theory.
Environment fit moderates velocity. Increasing agentic capacity should produce more durable throughput when the relevant engineered environment fits the work, and more churn, escaped failure, or repeated intervention when it does not. The effect of velocity therefore depends on whether the environment provides the representations, evidence, authority, and coherence the work requires.
Representation creates leverage. A task-relevant, trustworthy model should reduce the amount of lower-level state a human or agent must reconstruct to answer the engineering question at hand. The advantage should matter most as the required reasoning state grows. The claim fails if models impose maintenance and interpretation costs without extending reasoning reach or preserving quality.
Engineering capital amortizes judgment. Where useful engineering knowledge or recurring judgment becomes durable structure, later work over that surface should require less reconstruction, repeated judgment, or rework, and should carry less risk. The prediction includes its own boundary: the return lasts only while the asset remains fit. Stale models, obsolete validators, conflicting controls, and procedures that outlive their problem should lose value and can eventually impose net cost.
Each claim can fail. Environment quality may have little moderating effect; explicit models may add cost without extending reasoning reach; accumulated structure may cost as much as the judgment it was meant to retire. These are the theory's core propositions. Section 6.3 derives more specific predictions from them.
The reasoning-horizon proposition
The Representation-Leverage prediction rests on the scale problem introduced in Part I.
Reasoning-Horizon Proposition. Large software systems contain more potentially relevant state than any finite reasoner can keep active at once. Task-relevant models can extend the effective reasoning horizon by replacing implementation detail with semantically richer representations of the properties under consideration. The gain holds only while the representation is relevant, sufficiently faithful to the relation it claims, and cheaper to use than reconstructing the same knowledge from lower-level artifacts.
Agentic systems do not create this scale problem. Human engineers have always reasoned about large systems through abstractions because no engineer can keep the implementation of a Linux-scale system in active view. Commodity intelligence changes the economics and frequency with which the problem arises: larger delegated tasks and greater implementation volume increase the return on representations that keep relevant state available across reasoning episodes.
The engineering-capital proposition
The amortization prediction depends on what durable structure gives later work.
Engineering-Capital Proposition. Durable engineering structures constitute capital when future work inherits useful capacity from them: less reconstruction, less repeated judgment, earlier or stronger evidence, safer action, or cheaper recovery. Models, mechanisms, architectures, procedures, and other engineering assets can all qualify. Their returns are local to the surfaces where later work inherits that capacity, and last only while the assets remain fit.
The analogy to technical debt is intentional. Debt makes future change more expensive; capital makes future change more productive. Neither status is permanent. Capital depreciates as systems, obligations, and tools change. A validator can become noise, a model can drift, and a once-useful constraint can begin blocking legitimate work. Governance therefore includes maintenance, reconciliation, and retirement. Accumulation is not the objective; productive capacity is.
Tolerance suggests another way to understand that productive capacity: engineering margin. Where an obligation supplies a meaningful boundary, margin is the room between the realized system and that boundary. Engineering capital can preserve or enlarge such room—for example through better isolation, stronger architecture, additional capacity, or machinery that prevents acceptable variation from drifting toward failure. Technical debt can consume it: an expedient realization may satisfy today's obligations while leaving later work less room to change safely. Some debt raises the cost of change without moving any measurable property toward violation, and many qualitative software obligations admit no useful scalar distance from their boundary. MAGE therefore does not equate technical debt with lost margin, nor does it provide a general measure of engineering margin.
6.1.5 Modeling Expands the Surface of Alignment
Modeling and Alignment are distinct engineering activities. Modeling externalizes consequential knowledge and intent so that a reasoner can work from it; Alignment gives selected obligations authority over realization. Every meaningful Alignment mechanism embodies some conception of acceptable realization, but that conception need not exist as a separately represented model. A type can rule out an invalid state, a sandbox can deny a capability, a permission boundary can reject an action, and a test can block a regression while carrying its obligation locally.
Modeling matters because it changes which properties become tractable, and at what scale. Suppose the architecture requires every permitted route to a remediation service to cross quota authorization. Reasoning from implementation requires reconstructing routes from handlers, middleware, configuration, and dependency injection. Represent the permitted routes explicitly and the same architectural obligation becomes a graph question: does every allowed path cross the quota boundary? The property did not become less semantic. The representation changed the object over which the question is decided.
This clarifies the relationship between Modeling and Alignment. Alignment can operate directly over actions and artifacts. Modeling makes additional system-level properties available to Alignment. A model can also earn value without binding anything: if it saves later engineers and agents from repeatedly reconstructing the same system knowledge, it is already producing engineering capital. Authority adds another return when a stable obligation expressed over that representation can be enforced by the environment.
Repeated judgment can become machinery by two routes. If an obligation is already decidable over the current artifact, Alignment can encode it directly in a test, validator, constraint, permission, or gate. If the obligation requires repeated semantic reconstruction from implementation, Modeling may first change the object of reasoning by externalizing the relevant state or relation.
Call the boundary the determinization frontier: judgment on one side must still be supplied per instance; judgment on the other can be carried repeatably by the environment. Alignment moves the frontier when a property is already decidable. Modeling moves it when a better representation makes the property decidable. Figure 6.1-2 draws the two routes.
The frontier can move incrementally: from prose judgment to an explicit invariant; from an invariant to exhaustive checking over a bounded state model; or, for temporal claims, to a specification over executions.
The determinization frontier concerns obligations: which judgments can the environment carry repeatably? Tolerance characterizes the variation those obligations permit; degrees of freedom concern the realization choices engineering deliberately leaves open. Moving the frontier does not require eliminating those choices. A system can make consequential obligations and their tolerances increasingly explicit and authoritative while preserving many acceptable realizations.
Determinization need not eliminate freedom
Not every unmodeled choice is a degree of freedom. Some unmodeled regions contain obligations whose tolerances are still unknown, tacit, or awaiting a useful representation. A degree of freedom is what engineering has genuinely left open after the governing obligations are accounted for.
Degrees of freedom can also carry exploratory value. An engineer may leave a region open while the relevant distinctions are still being discovered; specifying it early can encode an ontology before experience shows which choices matter. Modeling may begin when obligations stabilize, or earlier, when the region grows too complex to explore without structure. In the second case, the model becomes part of the inquiry: it makes variation and interaction tractable, so consequential choices can be distinguished from genuinely free ones. Authority should follow that discovery selectively rather than attach to everything the model represents. The computation-graph episode in Part V illustrates this transition.
Degrees of freedom are therefore relative to a realization problem, not a fixed stock of choices in the product. A change can select among existing freedoms, narrow them by establishing a new obligation, or extend the realization surface and introduce new ones.
At the limit, a sufficiently complete realization model leaves little consequential freedom: realization becomes a transformation problem, and deterministic generation may be preferable to an agent. This limiting case connects MAGE to conventional model-based engineering; Section 6.5 returns to the relationship. Most software systems occupy a different point in the design space. Engineering specifies what must be true; many implementation choices remain acceptable. Commodity intelligence makes it economical to let autonomous realization choose among them.
Part V supplied complementary evidence: longitudinal depth from the originating case and comparative variation from independent industrial reconstructions. Neither supplies population-level effect sizes or causal estimates. The theory above is the proposed explanation of those observations; Section 6.3 states predictions that require stronger external tests.
MAGE predicts that abundant implementation makes environment fit increasingly consequential. Useful structure can enter ex ante from known engineering knowledge or ex post from experience. Its value depends on whether later work inherits useful capacity from it, and that value can depreciate.
The next question is where this account should apply.
Works Cited
- Conant, Roger C., and W. Ross Ashby. “Every Good Regulator of a System Must Be a Model of That System.” International Journal of Systems Science 1, no. 2 (1970): 89–97. https://doi.org/10.1080/00207727008920220.
- LinearB. “The Engineering Productivity Gap: How Elite AI Teams Are Pulling Away from the Rest.” LinearB, 2026. https://linearb.io/resources/ai-engineering-productivity-gap.
- Riggins, Jennifer. “AI Coding Got Faster. Why Didn't Engineering?.” The New Stack, August 9, 2026. https://thenewstack.io/ai-productivity-measurement-gap/.