§6.1 A Theory of MAGE
Chapter 5 ended with two views of agentic engineering: one system studied deeply enough to reconstruct mechanisms and sequence, and independent systems studied broadly enough to expose variation. This section proposes an explanation for both.
The theory is directional rather than numerical. It names the main parts, how they interact, and the relationships that should hold under stated conditions. §6.2 turns those claims into testable predictions.
The central claim is simple. Agentic capacity acts through the engineered environment that surrounds it. A strong environment turns more implementation capacity into durable progress; a weak or incoherent one turns the same capacity into churn, escaped defects, repeated intervention, and control friction. Those outcomes also change what later work inherits. Poor realizations can degrade architecture, representations, or conventions, making later work harder. Engineers can instead diagnose recurring problems and change the environment so that later work inherits better structure. As the environment improves, larger tasks and greater autonomy become feasible, exposing the next frontier.
Theory of MAGE — informal
Agentic capacity does not by itself produce engineering progress. It magnifies the fit—or misfit—between the work attempted and the engineered environment that surrounds it.
"Engineering progress" is intentionally domain-specific language: the theory is stated in engineering quantities because software engineering is the domain we observed and the source of its evidence. What the theory proposes transfers is the causal structure, not the software vocabulary. Agentic capacity acts through an environment, and the fit between that environment and the work governs how much raw capacity becomes useful output, how much failure escapes, and how much human attention the work consumes. Software gives these quantities concrete forms—landed changes, defects, review, rework—but the theory does not require source code as its output. §7.3 asks how far that structure carries beyond software.
Chapter 1 began from two limitations of autonomous realization. Reasoners have finite reasoning horizons, and consequential judgments entrusted to fallible inference retain residual risk. Chapters 2 and 3 supplied two engineering responses: change the representation over which reasoning occurs, and give selected obligations independent enforcement. Chapter 5 showed a broader set of responses in practice: improve the reliability of probabilistic judgment, catch bad realizations through independent evidence and admission, or move selected recurring judgments into durable structure.
Chapter 5 identified the economic tradeoff: engineering judgment can be exercised repeatedly during realization or carried forward in representations, controls, and other production structures. That chapter examined both responses as features of a production system. Process design changes how work, information, tools, authority, evidence, and admission are arranged around the fabricator. Tolerance determines which variation in the resulting product is acceptable and how the factory establishes that selected consequential properties remain within bounds.
This chapter develops a theory of that production system. We begin with deliberately narrow models of probability and cost. They expose repeated risk, retries, and the cost of reaching an acceptable realization while holding aside the full lifecycle of defect detection, recovery, human attention, representation maintenance, friction from controls, and changes to the environment. The dynamic model then widens the unit. The theory asks how agentic capacity and the engineered environment together shape engineering outcomes, how particular interventions change the probability and cost of acceptable realization, how independently evaluated obligations affect what may be admitted, and when durable engineering structure is worth its carrying cost.
6.1.1 The Dynamic Model
Figure 6.1-1 puts the theory on one diagram. Agentic capacity supplies model capability, concurrency, and raw implementation rate. The governed engineering environment supplies the representations, constraints, evidence, and enforcement through which that capacity acts. Together they shape three outcomes: durable throughput, defect escape, and human attention consumed per unit of useful output.
When the environment does not fit the work, the mismatch becomes visible as churn, recurring failures, escaped defects, repeated intervention, or friction from the controls themselves. These symptoms tell engineers that something is wrong. They do not say what is wrong: an engineer or organization still has to decide whether the signal is a local defect, a missing representation, a weak boundary, an inadequate control, or an obsolete piece of machinery. That diagnosis drives the change to the environment, which shapes the next round of work.
The process has a simple feedback structure. Prior engineering knowledge gives the environment a starting point. Work then exposes gaps between what the task requires and what the environment supports. Engineers diagnose those gaps and improve the environment. A stronger environment supports more ambitious work, which eventually exposes new gaps. The cycle repeats as the system, the available agentic capability, or engineering ambition changes.
This feedback structure has a precedent in cybernetics, which studied how regulators use information about a system to keep its behavior within acceptable bounds. Conant and Ashby's good regulator theorem makes representation central: under its assumptions, the simplest optimal regulator must embody a model of the system it regulates.11. Roger C. Conant and W. Ross Ashby, “Every Good Regulator of a System Must Be a Model of That System,” International Journal of Systems Science 1, no. 2 (1970): 89–97, https://doi.org/10.1080/00207727008920220.
MAGE applies a related idea to software engineering. The environment represents consequential properties of the system and its engineering process, observes what happens during the work, and changes when its representations or controls prove inadequate. The recurring question is simple: what must the environment know in order to govern the work?
Not all useful structure is discovered through failure. Existing architecture, standards, requirements, operating history, and lessons imported from earlier systems can seed the environment before an agent acts. Other needs become visible only in use. MAGE therefore has two routes to durable structure: before failure, when engineers deliberately encode a known obligation or useful representation, and after experience, when use exposes a recurring cost or risk worth turning into shared structure. Governance conversion names the latter route; it is not the origin story of every model or mechanism.
Two conditions determine whether adaptation reaches the shared environment. First, someone has to recognize that pressure is structural rather than local. Second, that diagnosis needs a path to the shared environment—the architecture, models, mechanisms, or infrastructure that can prevent the failure class from recurring. Without diagnosis, pressure produces motion without direction. Without authority to change the shared environment, the lesson dies as a local patch. §6.3 treats both as scope conditions.
6.1.2 Where Engineering Can Apply Leverage
The dynamic model suggests a broader question: where can engineering intervene to improve agentic work? Engineers can improve the reasoning model itself; change the context, tools, memory, or harness through which it works; change the process by which work is decomposed, coordinated, reviewed, and approved; change the representations over which consequential reasoning occurs; or change the evidence and controls that determine which results acquire consequence. These interventions can substitute for or complement one another. A stronger reasoner may need less support from its environment; a better representation may make the same task tractable for a weaker reasoner; a stronger admission mechanism may tolerate a less reliable producer.
Those intervention points are the constructs the rest of this chapter models. Table 6.1-1 reads each of them back onto the factory Chapter 5 described.
| Modeled construct | Meaning in the theory | Factory interpretation | Principal effect |
|---|---|---|---|
| — task | Engineering work to be realized | The controlled change entering the factory | Determines what must be produced and which obligations apply |
| — reasoning model | Capability of the reasoner performing delegated work | The capability of the fabricator | Changes how reliably and cheaply intent can be interpreted and realized |
| — representation | Engineering information made available for reasoning | A way of carrying engineering knowledge and selected product tolerances | Changes what must be reconstructed through inference and what can be reasoned about explicitly |
| — harness and process | Context, tools, decomposition, routing, coordination, and other organization of work | Process design around the fabricator | Changes the conditions under which interpretation and realization occur |
| — Alignment strategy | Independent evidence and enforcement for selected obligations | Measurement and admission machinery associated with selected tolerances | Changes whether unacceptable realizations acquire consequence |
| Probability that an attempt produces an acceptable candidate | Reliability of production under the chosen fabricator, representation, and process | Determines retries and expected realization cost | |
| Probability that an unacceptable candidate escapes governing mechanisms | Residual failure of evidence and admission | Determines residual assurance after production | |
| Cost of an attempt | Cost of operating the factory for one realization attempt | Captures inference, tools, realization, validation, and human intervention as appropriate | |
| Cost of retaining useful engineering structure | Cost of keeping models, policies, skills, tests, generated configuration, controls, and correspondence machinery useful as the software changes | Determines whether durable structure is cheaper than repeated reconstruction | |
| Governed engineering environment inherited at step | The factory as changed by prior work and experience | Captures path dependence and governance conversion |
One emerging school of thought puts substantial leverage in process design around the agents. Hassan's Agentic Software Engineering treats trustworthy agentic development as an end-to-end system spanning people, process, tools, and artifacts. Yegge's Gas Town approaches the problem from orchestration: persistent work state, specialized agent roles, handoffs, coordination, supervision, and merge machinery organize many fallible agents into a functioning software factory. Both make the organization of production an important control surface.22. Ahmed E. Hassan, Agentic Software Engineering: Building Trustworthy Software with Stochastic Teammates at Unprecedented Scale, 1st ed. (2026), https://agenticse-book.github.io/.33. Steve Yegge, “Welcome to Gas Town,” January 1, 2026, https://steve-yegge.medium.com/welcome-to-gas-town-4f25ee16dd04.
MAGE places additional weight on the representations and enforceable obligations that work inherits from the engineered environment. Modeling changes the objects over which humans and agents reason; Alignment gives selected obligations evidence and enforcement independent of the producing reasoner; governance conversion lets later work inherit lessons learned by earlier work. These mechanisms complement process design. The theory therefore treats model capability, representation, harness and process, and enforcement as distinct variables whose effects may substitute for or compound one another; §6.4 returns to that design space.
6.1.3 What Makes an Environment Effective
More machinery does not mean a better environment. Ten thousand lines of stale checks are worse than one hundred lines that govern the obligations that matter. What counts is not how much machinery exists, but how well the environment fits the work.
An effective environment needs four things:
- Useful representations. The information needed for the work is explicit at an appropriate level of abstraction.
- Adequate coverage. Important obligations have appropriate evidence and enforcement.
- Coherence. Models and mechanisms reinforce one another rather than contradicting or needlessly duplicating.
- Economy. Their future value justifies their maintenance, friction, and context cost.
Together these determine how much consequential work still depends on fallible per-instance inference: the probabilistic surface of Chapter 1. Representation makes consequential knowledge explicit and, where its structure permits, easier to reason about. Coverage supplies evidence and enforcement. Coherence asks whether the parts work together. Economy asks whether avoiding repeated judgment is worth the cost of carrying the machinery.
Coverage does not require mechanizing every obligation. Some belong in a compiler rule or admission gate; others remain statistical, semantic, or consequential enough to require human judgment. A strong environment does not mechanize everything. It makes clear where judgment belongs and where paying for the same judgment repeatedly would be wasteful.
These are not stages of a maturity model.44. Terry B. Bollinger and Clement L. McGowan, “A Critical Look at Software Capability Evaluations,” IEEE Software 8, no. 4 (1991): 25–41, https://doi.org/10.1109/52.300034. Improve the part of the environment that is limiting the work. If engineers repeatedly reconstruct important state, improve the representation. If failures escape, strengthen evidence or enforcement. If controls have become a tax, reconcile or remove them.
6.1.4 Outcomes and Observables
A theory of agentic engineering should distinguish activity from progress. MAGE therefore treats raw implementation velocity as an input and tracks three outcomes: durable throughput, defect escape, and human-attention burden. Table 6.1-2 defines each outcome and identifies candidate observables.
| Outcome | What it measures | Candidate observables |
|---|---|---|
| Durable throughput | Raw implementation converted into lasting system progress | landed changes net of rework; churn share; the slope of sustained throughput |
| Defect escape | Relevant failures reaching later lifecycle stages or production | escaped-defect rate; reopen rate; silent-corruption escape |
| Human-attention burden | Human judgment required per unit of durable output | interventions per landed change; review time; conversion effort |
Durable throughput is useful system progress that survives rework, integration, and operation rather than merely landing in the repository. Defect escape records consequential failures that reach later lifecycle stages or production. Human-attention burden measures the scarce engineering judgment consumed per unit of durable progress.
HUMAN FACTORSInset — Human Attention Is Not a Free Control
Moving work from human execution to machine execution does not eliminate the cost of human attention if people must inspect the resulting stream. Human-factors research has studied this problem for decades. Wolfe and colleagues found that rare targets in visual search were disproportionately missed as target prevalence fell.55. Jeremy M. Wolfe et al., “Rare Items Often Missed in Visual Searches,” Nature 435 (2005): 439–40. Warm, Parasuraman, and Matthews synthesize a broader vigilance literature showing that sustained monitoring is not passive work: it consumes attentional resources, imposes substantial workload, and can produce performance decrements and stress.66. Joel S. Warm et al., “Vigilance Requires Hard Mental Work and Is Stressful,” Human Factors 50, no. 3 (2008): 433–41, https://doi.org/10.1518/001872008X312152. These findings apply particularly to human-machine systems in which people monitor automated processes.
Operational practice provides the corresponding field example. The Transportation Security Administration uses covert tests of real airport screening operations to identify vulnerabilities in passenger and checked-baggage screening. Government Accountability Office reviews document both the use of these tests and failures that prompted changes to testing and remediation.77. U.S. Government Accountability Office, Aviation Security: TSA Improved Covert Testing but Needs to Conduct More Risk-Informed Tests and Address Vulnerabilities, technical report GAO‑19‑374 (Washington, DC, 2019). The setting differs from software review, but the engineering lesson is general: consequential inspection performed by trained professionals remains a capacity-limited control. More automation upstream does not create more human attention downstream.
This matters directly for agentic engineering. If implementation output grows while assurance continues to require proportional human inspection, human attention becomes a throughput and assurance bottleneck. MAGE therefore does not seek to remove humans from consequential judgment. It seeks to concentrate human judgment where judgment is actually required. Recurring obligations that can be represented, evidenced, or enforced adequately should increasingly be carried by the engineered environment; scarce human attention can then be spent on ambiguous evidence, novel failures, tradeoffs, exceptions, and decisions whose consequences genuinely require expert judgment.
Human attention is therefore not merely an implementation detail of the governance process. It is one of the scarce resources the engineered environment is intended to conserve.
Several easy-to-count quantities measure activity or investment rather than outcome: commits, lines changed, control files, model-sync events, and support ratio. They can help explain what an environment is doing; they do not establish that it is producing a return.** The activity-versus-outcome gap already shows up in industry data. A 2026 benchmark of 83,000 developers across 253 organizations88. LinearB, “The Engineering Productivity Gap: How Elite AI Teams Are Pulling Away from the Rest,” LinearB, 2026, https://linearb.io/resources/ai-engineering-productivity-gap. reports large merge-rate gains among heavy AI users, yet finds yield falling among the heaviest even as merge rates roughly doubled; an industry analyst reads the same period as rising AI spend without proportionate shipping velocity, the difficulty having moved from comprehension to confidence.99. Jennifer Riggins, “AI Coding Got Faster. Why Didn't Engineering?,” The New Stack, August 9, 2026, https://thenewstack.io/ai-productivity-measurement-gap/. Activity is not the same as durable progress.
6.1.5 Three Core Predictions
Three predictions form the core of the theory. Each describes a different way the engineered environment can help.
First, the environment can make additional agentic capacity more productive rather than merely faster. Second, it can give a reasoner a better surface over which to work, reducing how much of the system must be reconstructed for each task. Third, it can preserve useful engineering judgment so that later work inherits it rather than purchasing the same reasoning again.
These effects are related, but they are not the same. Environment fit concerns whether capacity can be productively absorbed. Representation leverage concerns what the reasoner must reconstruct. Engineering capital concerns what later work gets to inherit.
Environment fit moderates the effect of velocity. Increasing agentic capacity should produce more durable throughput when the relevant engineered environment fits the work, and more churn, escaped failure, or repeated intervention when it does not. In either case, raw implementation velocity may increase; what changes is how much of that activity becomes durable progress. The effect of added velocity therefore depends on whether the environment supplies the representations, evidence, enforcement, and coherence the work requires.
Representation creates leverage. A task-relevant, trustworthy model should reduce the expected cost of reaching an acceptable realization, or increase the scale of work completed at matched quality, by reducing reconstruction and exposing consequential distinctions at a more tractable abstraction. The advantage should matter most as the required reasoning state grows. The claim fails if models impose maintenance and interpretation costs without extending reasoning reach or preserving quality.
Engineering capital amortizes judgment. Where useful engineering knowledge or recurring judgment becomes durable structure, later work over that surface should require less reconstruction, repeated judgment, or rework, and should carry less risk. The prediction includes its own boundary: the return lasts only while the asset remains fit. Stale models, obsolete validators, conflicting controls, and procedures that outlive their problem should lose value and can eventually impose net cost.
Each claim can fail. The environment may have little moderating effect; explicit models may add cost without extending reasoning reach; accumulated structure may cost as much as the judgment it was meant to retire. These are the theory's core claims. §6.2 derives more specific predictions from them.
The reasoning-horizon proposition
Why should representation create leverage in the first place?
Consider what happens when an engineer enters an unfamiliar system to answer a seemingly simple question: Can this service depend on that one? The answer may be distributed across source files, configuration, dependency injection, deployment descriptors, documentation, and conventions that exist only in the history of the system. Before reasoning about the dependency, the engineer must first reconstruct enough of the system to know what the dependency means.
An agent faces the same problem. A larger context window can postpone the limit, and better retrieval can bring useful evidence closer, but neither changes the underlying fact that implementation contains far more detail than most engineering questions require. DocAble encountered the limit literally: one document produced roughly 590,000 tokens of reasoning representation against a 272,000-token context. That result did not argue for a larger context window; it made windowing and whole-document compression the next representation problem.
That is the general move: a useful model changes the problem by presenting the consequential structure directly. This is the basis for the reasoning-horizon proposition.
Reasoning-Horizon Proposition. Large software systems contain more potentially relevant state than any finite reasoner can keep active at once. Task-relevant models can extend the effective reasoning horizon by replacing implementation detail with semantically richer representations of the properties under consideration. The gain holds only while the representation is relevant, sufficiently faithful to the relation it claims, and cheaper to use than reconstructing the same knowledge from lower-level artifacts.
General form. Consequential work can require more potentially relevant state than a finite reasoner can reliably keep active at once. Task-relevant representations can extend the effective reasoning horizon when they preserve the distinctions needed for the work while suppressing irrelevant detail. The software proposition is one specialization of this claim. Whether other domains admit representations with the same leverage is an empirical question.
A useful representation can shorten the reasoning path, not merely make each reasoning step easier. Without an explicit model, an agent may first have to reconstruct the relevant architecture, state, dependencies, obligations, or other relationships from lower-level artifacts before it can reason about the engineering question itself. When the environment already supplies a trustworthy, task-relevant model, the agent can skip much of that reconstruction and reason directly over the consequential structure. In this sense, the environment preserves some prior reasoning instead of forcing each attempt to reconstruct it.
The effect is analogous to using a map rather than repeatedly reconstructing a city from street-level observations. The map does not contain everything in the city, nor should it. Its value comes from preserving the relationships needed for a particular class of questions while suppressing detail that would otherwise have to be rediscovered. An engineering model provides the same kind of leverage when it preserves the distinctions relevant to the work.
Agentic systems do not create this scale problem. Human engineers have always used abstractions because no engineer can keep the implementation of a Linux-scale system in active view. Commodity intelligence changes the economics and frequency of the problem: larger delegated tasks and greater implementation volume increase the return on representations that preserve useful reasoning across episodes.
These representations do more than store facts. State machines expose transitions, dependency graphs expose relationships, contracts expose obligations, and assurance arguments relate claims to evidence. They put consequential structure into a form the reasoner can use directly. Chapter 2 offered six useful classes of such model, not a proof that they are necessary, sufficient, or minimal. Because representation creates leverage and can expand what Alignment can govern, choosing, reducing, and even learning representations is itself an engineering problem. The principle is simple: a representation is worthwhile when the reasoning it saves is worth more than the cost of building and carrying it.
Representation leverage can also work in the other direction. Chapter 4 showed the practical version of this move: use implementation signals and commodity intelligence to help recover latent models from an existing system. Repeated structures can reveal a domain concept the implementation already embodies but has never named; bottom-up analysis can recover an abstraction that later humans, agents, and mechanisms use directly. The representation is new as an explicit engineering artifact, even when the structure it captures was already latent in the system.
The stronger possibility is representation innovation. A reasoner need not be limited to recovering an abstraction that engineers already know how to name. By working across implementations, histories, measurements, failures, and existing models, commodity intelligence may propose a different decomposition, relation, state space, or other representation that makes an engineering question tractable. The innovation is not that the machine draws a novel diagram. It is that the representation exposes consequential structure that existing representations did not make practically available for reasoning or analysis.
This extends the reasoning-horizon proposition in both directions. Human-designed representations can compress engineering knowledge into forms that let machines reason beyond implementation detail. Machine-induced representations can make latent structure explicit. Machine-innovated representations may go further, exposing relationships that extend what humans and later machines can practically reason about.
Discovery does not confer trust. Neither induction nor innovation makes a representation trustworthy. Like any model in Chapter 2, a machine-proposed representation claims a relationship to the system and must earn the trust placed in that claim. A recovered model can be checked against the implementation it purports to summarize; a genuinely novel abstraction poses the harder question of what evidence establishes that it preserves the distinctions required by its intended engineering use. The stronger the reasoning, assurance, or enforcement built on the representation, the stronger that evidence must be.
Representation engineering can therefore work in both directions. Humans can construct abstractions that extend machine reasoning; machines can recover abstractions implicit in human-built systems; and machines may eventually help invent abstractions that extend the effective reasoning horizon of both. The research agenda returns to how representations can be selected, induced, invented, validated, learned, and reduced.
The engineering-capital proposition
Representation leverage explains how structure can make one episode of reasoning easier. The next question is what happens across episodes.
If engineers repeatedly reconstruct the same architectural fact, rediscover the same failure mode, or make the same review judgment, the organization repeatedly pays for essentially the same reasoning. But some judgments can be preserved. A dependency rule can become a validator. An architectural distinction can become a model. An operational lesson can become a procedure. Once that happens, later work begins from a different starting point.
This is the sense in which MAGE uses the term engineering capital: prior engineering work leaves behind productive capacity that future work can inherit.
Engineering-Capital Proposition. Durable engineering structures become capital when future work inherits useful capacity from them: less reconstruction, less repeated judgment, earlier or stronger evidence, safer action, or cheaper recovery. Models, mechanisms, architectures, procedures, and other engineering assets can all qualify. Their returns are local to the surfaces where later work inherits that capacity, and last only while the assets remain fit.
The metaphor is capital rather than memory. A document that merely records an old decision may preserve information; an engineering asset becomes capital when it changes the cost, capability, or risk of future work. Like other capital, it must produce a return to justify carrying it, and it can depreciate when the environment changes. What counts as a justifying return is not fixed either: as agentic capacity lowers the cost of constructing and maintaining such assets, investments that were once uneconomical can clear the threshold.
That inheritance need not depend on the continued presence of the engineer who supplied the original judgment. When an architectural distinction becomes a trustworthy model, a review judgment becomes a validator, or an operational lesson becomes a reliable procedure, some of the originating engineer's productive capacity now lives in the environment. A later engineer or agent can benefit from the result without reconstructing the reasoning that produced it.
The transfer is necessarily partial. Durable structure carries only the judgment it successfully represents or mechanizes. When requirements change, a new failure appears, or the representation itself becomes inadequate, engineering judgment is needed again. Engineering capital therefore does not eliminate expertise; it changes how often the organization must pay again for expert judgment on the same question.
The analogy to technical debt is intentional. Debt makes future change more expensive; capital makes future change more productive. Neither status is permanent. Capital depreciates as systems, obligations, and tools change. A validator can become noise, a model can drift, and a once-useful constraint can begin blocking legitimate work. Governance therefore includes maintenance, reconciliation, and retirement. Accumulation is not the objective; productive capacity is.
The analogy also connects MAGE to an emerging account of debt in AI-assisted software. Storey argues that technical debt in the implementation is only one of three interacting forms of debt. Cognitive debt accumulates when a team's shared understanding of the system erodes; intent debt accumulates when goals, constraints, and rationale are poorly externalized or maintained. Generative AI can accelerate all three by increasing implementation faster than teams can understand it or preserve why it exists.1010. Margaret-Anne Storey, “From Technical Debt to Cognitive and Intent Debt: Rethinking Software Health in the Age of AI,” Queue 24, no. 3 (2026), https://doi.org/10.1145/3807966.
A Governed Engineering Environment acts on each layer, although not in the same way. Architectural constraints, validators, models, and other durable structures can reduce technical debt by keeping later realizations inside selected structural boundaries. Explicit requirements, decisions, models, and obligations counter intent debt by preserving consequential knowledge outside any one person or reasoning episode. Cognitive debt is different: MAGE does not require every engineer to reconstruct every generated implementation in detail. It instead seeks to preserve the understanding needed for consequential engineering decisions in representations and evidence that humans and agents can inspect, challenge, and reuse.
This does not make the three debts disappear. A stale model can preserve obsolete intent; an incoherent collection of mechanisms can itself become technical burden; and no representation can substitute for understanding that was never captured in the first place. The point is narrower. Engineering capital gives selected knowledge, intent, and judgment a durable place to live, while governance determines which of those assets remain trustworthy enough for later work to inherit.
6.1.6 Modeling and Alignment Are Distinct—and Compound
The previous propositions explain why durable structure can help, but they combine two effects that MAGE keeps separate. An environment can help the reasoner reach a satisfactory answer, and it can independently control whether an unsatisfactory answer acquires consequence. MAGE calls the first Modeling and the second Alignment.
Modeling and Alignment are therefore distinct engineering activities. Modeling changes the representation through which engineers and agents reason; Alignment enforces selected obligations on realization. Either can produce value without the other.
A model can reduce reconstruction and make satisfactory realization easier without enforcing anything. Conversely, an Alignment mechanism can govern an obligation even when the producing agent never sees the representation or rule from which the mechanism decides. A type can exclude an invalid state, a sandbox can deny a capability, and a validator can reject a forbidden dependency even if the agent repeatedly proposes one.
Consider a simple architectural obligation: service A may not depend directly on service B.
One engineering problem is helping an agent understand the relevant architectural structure well enough not to propose the forbidden dependency. A useful structural model may make acceptable realizations easier to produce. A different engineering problem is ensuring that the forbidden dependency cannot enter the system even when the producer gets that relationship wrong. An independent admission check can provide that guarantee.
Cost and failure, separately
We can make the distinction precise by considering cost and failure separately.
Let denote the representation available to a reasoner. Let be the probability that one attempt produces a realization satisfying a particular obligation, and let be the cost of producing that attempt. A better representation can increase , decrease , or both. Under a simple retry model in which successive attempts have the same success probability and expected cost , the expected realization cost — the expected production cost required to obtain one acceptable realization — is
This is one role of Modeling: make acceptable realization cheaper to find. The expression is deliberately simple: later sections relax the fixed-attempt assumption when successive changes alter the engineering environment.
The ratio is useful because it exposes the interaction between cost and reliability. Halving the cost of an attempt and doubling its probability of success have the same first-order effect on this quantity. More generally, a more expensive attempt can be cheaper in expectation if its probability of satisfactory realization rises enough. A structured representation may therefore cost more to construct, maintain, or supply than an informal description and still reduce the expected realization cost. A representation is worth preferring to whenever
Conversely, cheap natural-language articulation can be expensive in expectation when consequential intent is incompletely expressed or must be repeatedly reconstructed by the reasoner. A better reasoning model can compensate for a weak representation by raising , but at greater inference cost; a better representation can instead change the problem the reasoner must solve. Modeling can therefore trade representational investment against both realization cost and reasoning capability — which is why commodity intelligence does not by itself retire the case for representation. §6.4 makes the trade concrete: a trustworthy dependency model can turn work that appeared to require a frontier model into a graph problem a cheaper reasoner handles reliably — a consequence of this inequality, not a separate observation.
Representation precedes model selection. Before asking which reasoner should solve a problem, ask what representation would make the problem available for reasoning. In DocAble this became a practical cost rule: improve the representation before buying more inference.
The probabilities in these expressions belong to a production system. Figure 6.1-2 puts them on the factory Chapter 5 asked the reader to inspect.
The distinction matters because improving the fabricator and improving the tolerance act differently on this system.
The aggregate probability hides several distinct opportunities for failure. Let denote correct encoding of consequential intent, correct interpretation of that encoding, and correct realization of the interpreted intent. For a fixed representation , let denote the probability that consequential intent is correctly encoded in that representation. Then
This is a chain decomposition, not an independence assumption: the stages will not in general be statistically independent. Its value is that the three stages fail for different reasons:
- Encoding — did the human correctly externalize consequential intent into the available representation?
- Interpretation — did the reasoner correctly recover that intent from the representation?
- Realization — did the reasoner successfully realize what it understood?
The first term is easy to overlook in an agentic setting. Representation is not valuable only because a machine may interpret a structured model more reliably than prose. It can also help the engineer express the intended property correctly in the first place. Different representations make different distinctions visible, force different decisions to be made explicitly, and make omissions easier or harder to notice. A representation can therefore improve even if it leaves unchanged. This is one reason Modeling is an engineering activity rather than merely a technique for supplying context to an agent.
Representation is not the only route to that term. Because a realization that misses the engineer's meaning is evidence about the encoding, the comprehension amplification of §4.1 can raise indirectly: capable agents return that evidence quickly enough for the engineer to encode the intent again.
The Chapter 5 distinction identifies where different factory interventions enter this decomposition. Process design changes the conditions under which interpretation and realization occur: context, tools, decomposition, routing, permissions, and other features of the harness can change , , their cost, or several at once. A better agent raises the same terms. Explicit tolerances act differently. They bear on — did we correctly represent the property we actually care about? — and they make selected consequential properties available for independent evaluation, so the factory need not rely solely on successful producer interpretation and realization to determine whether those properties hold. The realization can instead be measured against the represented tolerance and rejected.
The probability of correct interpretation also depends on how much inference the representation asks the agent to supply. That burden is not constant across engineering tasks. A capable model has encountered enormous numbers of ordinary database interactions, standard algorithms, common framework idioms, and familiar architectural patterns. When the intended solution follows such well-represented patterns, a sparse specification may still yield a high probability of correct interpretation. The agent can supply much of the missing detail from what it already knows.
The situation changes as the intended approach becomes less conventional. Novel algorithms, unusual architectural constraints, domain-specific invariants, and decisions that deliberately depart from common practice provide less familiar structure for the agent to recover. Other things equal, the more novel the intended engineering decision, the less confidence we should place in correct interpretation of what has been left implicit. Specification substitutes explicit engineering intent for inference from prior patterns.
But familiarity solves only the production problem. It does not establish the property. If an agent usually interprets "connect this service to the database" in the conventional way, that prior may be entirely adequate for a prototype. If the production system depends on a particular transaction boundary, consistency property, authorization rule, or failure behavior, relying on the same prior provides no independent basis for determining that the obligation was preserved. The agent may probably do the familiar thing; unless the consequential property is represented, the engineering environment cannot reliably distinguish following the intended pattern from merely producing something that looks right.
Familiarity is not assurance. This is why increasing does not eliminate the role of explicit models or Alignment. Learned priors can reduce how much must be specified for realization. They do not by themselves provide an assurance surface for consequential properties.
The same limit applies to the production conditions that raise that factor. Better context, decomposition, tools, examples, and review all improve how an agent interprets natural-language intent. None of them changes what happens next. If the consequential obligation itself remains in natural language, establishing that the result satisfies it may require another act of probabilistic interpretation. Ambiguity is only part of this. The evaluation inherits the fallibility of the production, so a factory that improves only its production conditions has strengthened one probabilistic step and left the other standing. Tolerance changes the structure of the problem by supplying an independently evaluable boundary around acceptable realizations.
Thus depends not only on the quality of the encoding and the capability of the agent, but also on the familiarity of the intended interpretation relative to what the agent already knows.
The first factor sits largely at the human/factory interface, not in machine reasoning. Two different failures hide there: the engineers may never have articulated a consequential distinction, and the reasoner may have failed to recover it. The entitlement failure from Chapter 2 is the worked instance. The scalar representation read each group's unlimited flag and still returned a single finite sum: its output type had no place for the distinction, so the distinction died before any reasoner saw it. No downstream capability could have recovered what the representation could not express.
A sufficiently capable reasoner cannot recover information that the engineering environment failed to express. Capability compensates for difficult interpretation; it does not reliably compensate for missing intent. Modeling can therefore improve realization economics in two ways before implementation begins: by helping humans state what they mean, and by giving machine reasoners a better surface from which to recover it. And Modeling sometimes goes beyond raising these probabilities at all — the determinization frontier later in this section names the case where a change of representation makes a reconstructed semantic property mechanically decidable, removing the per-instance judgment altogether.
The limit case makes the asymmetry exact. Write for the reasoning model. Within the three-stage model, for a fixed representation , as approaches a perfect reasoner and the interpretation and realization factors approach one,
Additional failure modes can only lower the overall probability of satisfactory realization. Even perfect interpretation and realization remain bounded by whether the engineering intent was correctly encoded in the representation available to the factory. The bound restates this book's account of engineering responsibility in economic terms: engineers can delegate increasing amounts of realization while remaining responsible for the result, and an organization that has not determined and externalized the obligations governing the work cannot discharge that responsibility by purchasing a more capable reasoner.
Perfect fabrication does not eliminate the engineering problem upstream of fabrication. It exposes it, quickly. A stronger reasoner can compensate for ambiguity by doing more inference: it can recover conventions, inspect surrounding artifacts, search for analogous cases, test hypotheses, and spend additional computation reconstructing what the engineer probably intended. Those capabilities raise the interpretation and realization factors. They cannot reliably recover consequential knowledge that never entered the factory.
This is not necessarily a defect. Agile practice exploits rapid fabrication precisely because working realizations can produce feedback sooner. The danger depends on whether incorrect upstream assumptions are exposed by that feedback or survive it as latent failures that surface only in production. Feedback of that kind does not contradict the bound. It exploits the coupling the dynamic model later in this section makes explicit: each realization changes the engineering environment the next one inherits.
Read in factory terms, the bound limits process design. As agent capability improves, process design can raise the probability that a fabricator correctly interprets and realizes engineering intent. It cannot make an incorrectly encoded obligation correct. Tolerance therefore exposes a different limit. Where a consequential property is represented explicitly and evaluated independently, improvements in fabrication drive interpretation and realization error downward while the remaining error approaches the probability that the engineering representation itself is wrong.
Nor does the bound argue for complete specification. Determining nearly every consequential and inconsequential realization decision in advance would recreate much of the economic problem that defeated the earlier software factories (Chapter 5). The new economic possibility is different: a capable fabricator can be given enormous freedom over realization while the engineering environment explicitly preserves the smaller set of decisions whose resolution matters.
Carrying cost completes the comparison. The inequality above compared realization costs alone; a representation must also be constructed and maintained. The total cost of holding a representation and realizing through it is
is broader than the cost of maintaining a formal model. Chapter 5 showed factories carrying engineering knowledge in models, instructions, policies, skills, tests, generated configuration, context systems, and other production structures. These structures differ in how much judgment they preserve and how strongly they govern realization, but all incur some cost to remain useful as the software and its environment change.
An explicit model typically raises this carrying cost. It can still be economical even when it is expensive to construct and each realization through it is individually more expensive: the model need only reduce expected realization, reconstruction, failure, or assurance cost enough over its useful life. Natural language is therefore a legitimate point in the design space, not a defect to be engineered away; for many tasks its carrying cost is extremely low.
Commodity intelligence changes both sides of this investment. Greater implementation capacity can justify richer engineering structure, while the same capacity can reduce the cost of carrying that structure. The carrying term is therefore not a fixed parameter of the environment: the same capacity that raises the volume of realization the environment must govern also helps construct and maintain the models, checks, and correspondence machinery that govern it.
The factorization also shows where reasoning capability can substitute for representation. When natural language leaves a difficult interpretation problem, one response is a durable representation; another is a more capable reasoner, more inference, richer context, or additional supervisory reasoning. Those interventions may raise and , but their cost is incurred again as production continues. For an inexpensive reasoner and a frontier reasoner over the same natural-language representation , it may well be that
A factory can therefore compensate for weak explicit representation by spending heavily on reasoning capability. That is not an objection to the theory; it is a prediction of it. One arrangement pays repeatedly in the interpretation factor at frontier inference prices; another pays once in carrying cost so that a cheaper reasoner suffices. Keeping consequential properties inside natural-language inference therefore makes reliability a recurring purchase: better models, more reasoning, more context, repeated review, stronger supervising agents, bought again on every change. Representing a property explicitly converts some of that recurring inference into durable engineering structure, moving its cost into the carrying term above, where repeated use amortizes it. The factory a vendor sells — capable fabricator, generic supervisory machinery, none of your consequential intent (Chapter 5) — thereby becomes a theoretically interesting object rather than a descriptive category: how much engineering structure can commodity reasoning capability economically replace?
This does not imply that every recurring judgment should become a model. The comparison depends on recurrence, carrying cost, inference cost, consequence, and how much the representation actually improves production or assurance.
Repeated work changes the economics again. For a single task, writing a structured model may be extravagant. If the same representation supports realizations, then, under a simple linear maintenance model that splits the carrying cost into its fixed and recurring parts,
Build once, carry repeatedly, realize repeatedly. The fixed cost amortizes across . More importantly, the judgment embodied in the representation is reused: one policy shapes thousands of changes; one learned constraint protects every later task. The factories that accumulate shared policy, models, context, and permissions as scale rises are not merely adding process. They are amortizing judgment.
§5.3 showed factories making different choices about where engineering knowledge lives and what their agents must reconstruct. The decomposition explains why those arrangements can coexist: a factory may spend on explicit representation, on reasoner capability, on human supervision, or on some combination, and each purchase moves different factors. The relevant question is not whether a factory uses models in the abstract, but what probability and cost of satisfactory realization its arrangement produces — and what engineering knowledge survives for the next task.
Now let denote an independent admission mechanism. In the factory vocabulary of Chapter 5, is part of the machinery by which a selected tolerance becomes consequential: evidence establishes whether the realization remains within the represented acceptable region, and admission determines what happens when it does not. In the strongest region of Alignment, suppose is a sound deterministic predicate for the covered obligation: an unacceptable candidate cannot pass. Then
That guarantee does not require the producer to reason well about the obligation. The agent can be shown no structural model at all, repeatedly propose forbidden dependencies, and still be prevented from merging them. This is one role of Alignment: make consequence independent of whether the producing reasoner got the covered judgment right.
But assurance without useful representation can be expensive. If each independently generated candidate satisfies the covered obligation with probability , each attempt costs to produce and to validate, and rejected candidates are regenerated until one passes, then
Alignment can therefore provide the same admission guarantee over two reasoning surfaces at very different realization costs.
Return to the architectural example. Suppose the rule is checked perfectly at merge time. Hiding the structural model from the agent does not weaken that gate: forbidden changes still cannot merge. But the agent may repeatedly reconstruct the relevant architectural relationship incorrectly and collide with the gate. Expose an appropriate structural model to the same agent and the guarantee need not change at all. What changes is the cost of reaching a candidate the gate will accept.
Alignment protects the boundary; Modeling reduces how often work collides with it.
The division can be stated by stage. Modeling can act at several points before admission: helping engineers state consequential intent, helping reasoners recover it, and making satisfactory realization easier. Alignment acts afterward or alongside realization, by determining whether a covered failure can acquire consequence.
The four combinations therefore have different engineering meanings (Table 6.1-3).
| Modeling ↓ / Alignment → | No Alignment | Alignment |
|---|---|---|
| No Modeling | Repeated reasoning from a poor surface; residual failure can escape | Covered failure cannot escape, but rejection and regeneration may be expensive |
| Modeling | Satisfactory realization becomes cheaper or more likely, but residual failure can still escape | Satisfactory realization becomes cheaper while covered failure remains independently blocked |
The table is schematic rather than exhaustive. Modeling may itself improve reliability, and Alignment mechanisms may be statistical rather than deterministic. The point is that their primary contributions are separable.
Software engineering changes the problem
The analysis so far uses a simplification to isolate repeated probabilistic exposure: if consequential judgments have some probability of success and successive exposures are treated as independent, their joint probability of success falls as the number of exposures grows. The retry model likewise holds the probability and cost of an attempt fixed while asking what additional attempts purchase. These models are useful because they isolate the effects of exposure and repetition.
Software engineering requires a different model because one realization changes the conditions inherited by the next. Winters et al. characterize software engineering as "programming integrated over time": software is developed, modified, and maintained as both the system and its environment change.1111. Titus Winters et al., Software Engineering at Google: Lessons Learned from Programming over Time (O'Reilly Media, 2020). One realization therefore becomes part of the architecture, conventions, representations, tests, and other structure inherited by subsequent work. A good change can make later work easier to reason about; a locally acceptable but structurally poor change can make it harder.
Creation and change. The factorization above describes whether one realization succeeds, whether that realization creates a system or changes an existing one. Software engineering makes the second case different because a change does not begin from an empty environment. It inherits architecture, representations, conventions, obligations, tests, and engineering capital produced by earlier work. The probability of satisfactory realization must therefore depend on the environment at the time of the change.
For this question, independence leaves out the relationship that matters. Let the probability of satisfactory realization at step be
where is the governed engineering environment inherited at that step. The resulting realization then helps determine the environment available to subsequent work:
- Creation — satisfactory realization from an initial engineering environment .
- Change — satisfactory realization given the inherited environment .
- Inheritance — realization changes the environment available to later work, .
The factory therefore produces a sequence of changes, not independent realizations.
This coupling creates path dependence. A locally acceptable change can increase architectural irregularity, obscure system structure, introduce inconsistent conventions, or otherwise make later work harder to reason about. Later agents then operate over a worse reasoning surface, potentially reducing the probability of satisfactory realization. Architectural decay can therefore become self-reinforcing: poor structure makes good realization harder, and subsequent realizations can degrade the structure further.
The reverse trajectory is possible as well. A realization that improves representations, architecture, constraints, or other engineering capital can make later work easier to reason about and govern. The environment then carries useful structure forward rather than forcing each subsequent reasoner to reconstruct it.
The independent model and the dynamic model therefore answer different questions. The first isolates the effect of repeated probabilistic exposure under fixed conditions. The second preserves the relationship between successive realizations and the environments they create. These linked trajectories are the mathematical counterpart of the feedback loops in Figure 6.1-1.
Modeling can also enable Alignment
The two principles interact in a second way. Some obligations can already be checked mechanically, and Alignment can govern them directly. Others are semantic only because the relevant state must be reconstructed from lower-level implementation. Suppose the architecture requires every route to a remediation service to cross quota authorization: reasoning from implementation means reconstructing routes from handlers, middleware, configuration, and dependency injection, but representing the permitted routes explicitly turns the obligation into a graph question—does every allowed path cross the quota boundary? The property did not become less consequential; Modeling changed the object over which it is decided, turning a repeated semantic judgment into a checkable property.
There is a broader pattern here. Some engineering questions are hard because the answer is intrinsically judgmental. Others are hard because the information needed to answer them is buried in the wrong representation. These cases look similar when an engineer or agent is staring at source code: both require reasoning. But they have very different engineering possibilities. If the underlying property is genuinely judgmental, better representation may help without eliminating the judgment. If the property becomes mechanically decidable once the relevant structure is made explicit, Modeling has done something stronger: it has moved a decision from repeated inference into repeatable machinery. The boundary between those cases matters enough to name.
Call the boundary between these cases the determinization frontier: judgment on one side must still be supplied per instance; judgment on the other can be carried repeatably by the environment. Alignment can mechanize an obligation that is already decidable. Modeling can move the frontier by changing the representation until a previously reconstructed property becomes decidable. Figure 6.1-3 draws the two routes. The frontier is not a wall around what agents may do. It separates judgments the environment must repeatedly purchase from judgments it can carry forward itself. Read in factory terms, that separation is partly economic. On one side sit the judgments a factory continues to buy probabilistically through its production process; on the other, the properties it has chosen to represent explicitly and carry as durable tolerances.
My default in DocAble was simple: everything that could be deterministic was. I would rather have a deterministic result I can rely on than a probabilistic result where the problem does not require judgment. Heading structure, file metadata, schema validity, and similar properties can be checked directly. Writing useful alternative text cannot. The determinization frontier separates those cases: use commodity intelligence where semantic judgment is required; use deterministic machinery where it is not.
Moving the frontier reduces the probabilistic surface: fewer consequential judgments depend on the agent getting them right. It does not make the agent deterministic, and it does not require eliminating implementation freedom. The environment can tightly govern what must be true while leaving the agent free to choose among many acceptable realizations.
The frontier can move incrementally: from prose judgment to an explicit invariant; from an invariant to exhaustive checking over a bounded state model; or, for temporal claims, to a specification over executions.
Moving the frontier does not mean closing off realization. The frontier concerns obligations: which judgments the environment can carry repeatably. The acceptable ways to satisfy them remain degrees of freedom for the agent to explore.
At the limit, a sufficiently complete realization model leaves little consequential freedom: realization becomes a transformation problem, and deterministic generation may be preferable to an agent. This limiting case connects MAGE to conventional model-based engineering; §7.2 returns to the relationship. Most software systems occupy a different point in the design space. Engineering specifies what must be true; many implementation choices remain acceptable. Commodity intelligence makes it economical to let autonomous realization choose among them.
The joint engineering problem
We can now put the pieces back together.
Modeling can make satisfactory realization easier to find. Alignment can prevent selected failures from acquiring consequence. Better environments can preserve useful judgment for later work, but every representation and mechanism also costs something to build, maintain, reconcile, and carry.
The objective is therefore not to maximize Modeling, maximize Alignment, or minimize the probabilistic surface at any cost. It is to engineer an economical division of labor among the reasoner, the environment, and the human judgment that remains.
For a task and its obligations, an engineer chooses both a representation and an Alignment strategy :
The expression is deliberately schematic. Modeling principally changes the cost and difficulty of satisfactory realization; Alignment principally changes the probability and consequence of unacceptable realization; both impose construction and carrying costs. This is where the factory's standing choice gets priced: the realization term absorbs inference quality bought again on every change, and the carrying term absorbs the structure that makes a property durable as a tolerance instead. In high-assurance settings the problem may instead be stated as minimizing total cost subject to an assurance requirement,
For an adequately implemented deterministic gate over a decidable obligation, can be zero for that covered property even while realization remains probabilistic.
6.1.7 The Factory's Value Proposition
The expression above prices one realization. A factory is an investment in productive capacity, so its economics cannot be evaluated from the cost of constructing it alone. A physical factory incurs a startup cost to establish production, and the investment is justified when the products it subsequently produces are worth both that cost and the marginal cost of continued production.
Software preserves this structure but changes what is produced. A software factory generally does not manufacture each copy its customers consume: once software has been realized, another copy may be nearly free, and even a hosted service can distribute the benefit of one engineering change across many users. The recurring output is the one §5.1 named — the continued transformation of a maintained software system: features, repairs, adaptations, migrations.
We can therefore separate factory construction from continued production. For a factory that has produced changes,
The cost of a change includes the human attention required to direct, coordinate, review, diagnose, and intervene; the machine intelligence used for realization and reasoning; computational infrastructure; validation and rework; and any additional engineering structure created in producing the change. captures the initial investment required to establish the factory, and the cost of keeping its productive capacity usable. The pair enters this expression on both sides: representations, controls, and evidence raise and while potentially lowering each and the expected failure loss.
The corresponding marginal question is especially important:
A successful factory may be expensive to construct and economical to use. The investment is justified not because its first product was inexpensive, but because the resulting productive capacity makes a sufficiently valuable stream of subsequent changes cheaper, faster, safer, or possible at all.
Two Margins, Not One
Software adds a distinction manufacturing does not have, and it should be stated explicitly: there are two margins here, not one. The production margin asks what it costs the factory to make the next acceptable change to the software. The replication margin asks what it costs, once that change exists, to provide its benefit to the next user, document, transaction, installation, or copy. For packaged software the second can approach zero. For a hosted service it is not literally zero — compute, storage, and support scale with use — but it can remain small relative to the engineering cost of changing the product.
The Idle Factory Still Depreciates
Carrying cost also divides, and the division is where the physical analogy breaks most instructively. A semiconductor fabrication plant sitting idle is still enormously expensive: buildings, equipment, utilities, and staffing continue whether or not anything is produced. An agentic software factory can idle at almost no physical cost. Repositories, models, configuration, tests, and governance mechanisms sit on disk; cloud services scale toward zero; agent subscriptions stop.
The first term can be extraordinarily low for software. The second need not be: dependencies move, models cease to correspond to the system, tests go stale, and organizational knowledge disappears. Software factories can be cheap to leave physically idle while their engineering capital continues to depreciate.
What the Changes Are Worth
Against these costs stands what the stream of changes is worth, and that side of the comparison is deliberately left abstract, because it cannot be specified universally. A product factory may value capabilities delivered to customers; a security factory may value risk retired; an accessibility factory may value documents made accessible and remediation effort avoided. These quantities are themselves models of value chosen for a particular engineering or organizational decision — the same claim Chapter 2 made about metrics generally. The theory does not define what an outcome is worth. It asks whether the outcomes produced justify the full cost of the production arrangement.
The factory produces changes to the software; the software produces outcomes for its users. A company may spend substantially to produce one change and then distribute its benefit across millions of transactions at comparatively small incremental cost, so factory productivity and product value meet at a boundary rather than being the same quantity:
MAGE primarily concerns the economics of the first arrow. The organization's business model determines how the resulting capability becomes value.
Correctness Is Not the Factory's Objective
Correctness, however, is still not the factory's final objective. It is one of the mechanisms through which production becomes valuable. The lifecycle expression above prices a factory as a startup investment plus a stream of change costs, a carrying cost, and an expected failure loss. The probability model explains those terms from the inside. Higher increases durable throughput, because more attempts satisfy their obligations, and lowers the cost of a change by avoiding failed attempts, repeated inference, rework, intervention, and recovery. Alignment exchanges validation cost for reduced escape risk — spending inside the change term to shrink the expected failure loss. Models and mechanisms preserve successful reasoning as engineering capital rather than requiring the factory to purchase it again on every realization.
But correctness alone does not establish value. A factory can reliably and efficiently produce software that nobody needs. The probability model explains how engineering intent survives fabrication; whether successfully realizing that intent produces outcomes worth their full cost is the value question the factory economics ask, and no correctness result answers it.
This provides the larger economic interpretation of Modeling and Alignment. They are not valuable because explicitness is intrinsically desirable, nor because every decision should be mechanized. They are valuable where the cost of externalizing, preserving, and enforcing consequential knowledge is lower than the expected cost of repeatedly reconstructing it, incorrectly realizing it, or allowing its violation to escape.
Commodity intelligence changes the terms of that comparison. As reasoning becomes cheaper and more capable, more decisions can economically remain with the fabricator; at the same time, greater realization capacity increases the amount of work that can cross a weakly governed boundary before a human can inspect it. The engineering problem therefore does not disappear as fabrication approaches perfection. It moves toward deciding what must be preserved, what may safely be inferred, what must be established before admission, and what the resulting production capability is worth.
Chapter 5 supplied two kinds of evidence: longitudinal depth from the originating case and comparative variation from independent industrial reconstructions. Neither supplies population-level effect sizes or causal estimates. The theory above is the proposed explanation of those observations. §6.2 states predictions that require stronger external tests; §6.3 asks where they should be expected to hold.
Works Cited
- Conant, Roger C., and W. Ross Ashby. “Every Good Regulator of a System Must Be a Model of That System.” International Journal of Systems Science 1, no. 2 (1970): 89–97. https://doi.org/10.1080/00207727008920220.
- Hassan, Ahmed E. Agentic Software Engineering: Building Trustworthy Software with Stochastic Teammates at Unprecedented Scale. 1st ed. 2026. https://agenticse-book.github.io/.
- Yegge, Steve. “Welcome to Gas Town.” January 1, 2026. https://steve-yegge.medium.com/welcome-to-gas-town-4f25ee16dd04.
- Bollinger, Terry B., and Clement L. McGowan. “A Critical Look at Software Capability Evaluations.” IEEE Software 8, no. 4 (1991): 25–41. https://doi.org/10.1109/52.300034.
- Wolfe, Jeremy M., Todd S. Horowitz, and Naomi M. Kenner. “Rare Items Often Missed in Visual Searches.” Nature 435 (2005): 439–40.
- Warm, Joel S., Raja Parasuraman, and Gerald Matthews. “Vigilance Requires Hard Mental Work and Is Stressful.” Human Factors 50, no. 3 (2008): 433–41. https://doi.org/10.1518/001872008X312152.
- U.S. Government Accountability Office. Aviation Security: TSA Improved Covert Testing but Needs to Conduct More Risk-Informed Tests and Address Vulnerabilities. Technical Report GAO‑19‑374. Washington, DC, 2019.
- LinearB. “The Engineering Productivity Gap: How Elite AI Teams Are Pulling Away from the Rest.” LinearB, 2026. https://linearb.io/resources/ai-engineering-productivity-gap.
- Riggins, Jennifer. “AI Coding Got Faster. Why Didn't Engineering?.” The New Stack, August 9, 2026. https://thenewstack.io/ai-productivity-measurement-gap/.
- Storey, Margaret-Anne. “From Technical Debt to Cognitive and Intent Debt: Rethinking Software Health in the Age of AI.” Queue 24, no. 3 (2026). https://doi.org/10.1145/3807966.
- Winters, Titus, Tom Manshreck, and Hyrum Wright. Software Engineering at Google: Lessons Learned from Programming over Time. O'Reilly Media, 2020.