§4.4 Engineering the Environment
The preceding sections watched the same thing happen from three directions. §4.1 showed models becoming governing constraints one episode at a time; §4.2 followed one problem until a family of models governed an architecture; §4.3 recovered structure a system already owned. Each time, the environment ended holding more than it started with: more representations, more mechanisms, more preserved judgment. This section asks about that destination directly. What should the engineering environment itself know, do, check, and preserve?
4.4.1 The GEE as a Distributed Cognitive System
Not every important judgment should become deterministic. There are two complementary ways to improve recurring engineering work. One is to improve the conditions under which a reasoner makes a judgment: give it better representations, domain knowledge, examples, procedures, tools, and context so that a satisfactory realization becomes more likely. The other is to enforce selected obligations in the environment so that satisfying them does not depend entirely on the reasoner choosing correctly.
The distinction is not between two kinds of engineering artifact, but between two effects those artifacts can provide. A model can guide an agent and state properties consumed by a validator. A skill can improve judgment and invoke controls that enforce selected obligations. Evidence can inform a decision or feed an admission gate. A governed engineering environment composes these roles rather than assigning each artifact permanently to one side.
Nor is the choice simply between probabilistic judgment and deterministic enforcement. Start with the obligation and ask what representation makes it tractable, then use the cheapest judge adequate to the representation and the decision at hand. Representation can change that allocation: a judgment that initially requires expensive interpretation may become mechanically decidable once the relevant states, relations, or invariants are explicit. The converse failure also exists. A cheap judge can be perfectly reliable over an inadequate representation and still answer the wrong engineering question. In DocAble, pixel variance was easy to measure but could not decide whether a low-variance image carried semantic content; the representation had discarded the distinction the obligation required. The engineering move was not to tune the threshold harder, but to route the unresolved semantic question to a richer judge.
Use the weakest representation that preserves the distinction the obligation requires, and the cheapest judge adequate to decide it. Then decide, separately, whether the resulting judgment warrants enforcement. A deterministic check may remain advisory; a probabilistic evaluator may sometimes be reliable enough to support enforcement. Adequacy of judgment and whether the judgment should be enforced are separate engineering decisions.
There is an older way to understand the system this composition produces. Work on distributed cognition studies reasoning performed by systems of people, representations, procedures, and artifacts rather than locating all cognition inside one individual.11. Edwin Hutchins, Cognition in the Wild (MIT Press, 1995). External representations matter in that account because they do cognitive work: they preserve state, change what operations are easy, and coordinate reasoning across participants. A Governed Engineering Environment has a similar systems character, although it adds an engineering concern that distributed cognition does not by itself supply: selected obligations can be enforced by mechanisms outside the immediate reasoner.
COGNITIVE SCIENCEInset — The GEE as a Distributed Cognitive System
Studies of distributed cognition ask us to move the boundary of analysis outward. Edwin Hutchins's classic study of ship navigation, for example, does not explain navigation solely by asking what one navigator knows. Position is established through a system of people with different roles, instruments, charts, procedures, communications, and intermediate representations. The relevant cognitive system is larger than any individual participating in it.1
The same perspective is useful for agentic engineering. A coding agent need not carry the whole engineering problem inside its context window. Models preserve selected facts about the system; skills preserve procedures; tools perform operations that need not be approximated through language-model reasoning; evidence records what occurred; people or agents supply judgments that remain difficult to evaluate mechanically. The capability of the resulting engineering system is therefore not simply the capability of its foundation model.
External representations matter for more than storage. Work on distributed cognitive tasks shows that the form of an external representation can change the operations required to solve a problem.22. Jiajie Zhang and Donald A. Norman, “Representations in Distributed Cognitive Tasks,” Cognitive Science 18, no. 1 (1994): 87–122, https://doi.org/10.1207/s15516709cog1801_3. A state machine exposes legal transitions; a dependency graph exposes reachability and coupling; a measurement model exposes a bound. The representation changes the reasoning problem presented to whoever—or whatever—reasons next.
A GEE adds another step. Its surrounding machinery can sometimes act on the represented property. A dependency model may help an agent understand the architecture, but a validator can also compare an observed dependency with an architectural obligation, and a gate can refuse the change. The environment therefore does not merely distribute cognition across reasoners and artifacts. For selected obligations, it can enforce what those representations express.
This perspective helps explain why replacing one model with a more capable one does not make the surrounding engineering environment irrelevant. The environment carries system-specific knowledge, procedures, evidence, and controls that no general-purpose reasoner should have to reconstruct for every task. Better reasoners change what that environment needs to supply. They do not eliminate the value of putting durable engineering knowledge in the system around them.
This is the system the chapter has been assembling all along. The kernel of §4.1, the model family of §4.2, and the recovered structure of §4.3 are not documentation attached to a system. They are working parts of a distributed cognitive system that includes the engineers and agents reasoning within it—with the additional MAGE property that, for selected obligations, the environment does not merely help the reasoning but enforces its conclusions.
The boundary between the two effects is not fixed. A judgment that initially resists enforcement may become better understood through repeated use, failure, and refinement. Its obligations may become clearer, suitable evidence may become available, or a qualitative distinction may become evaluable with sufficient reliability. At that point, some of what previously depended on judgment can be enforced by the environment; other portions may properly remain probabilistic. Figure 4.4-1 shows the recurring cycle. The return path is the governance conversion developed in §3.4: when an outcome or failure exposes missing representation, obligation, evidence, evaluation, or consequence, the organization can encode that missing structure so that future work inherits it.
The idea that an organization should learn by changing the structures that govern later action also predates MAGE. Organizational-learning research distinguishes correcting an observed error from changing the governing assumptions, rules, or norms that repeatedly produce action.33. Chris Argyris and Donald A. Schön, Organizational Learning: A Theory of Action Perspective (Addison-Wesley, 1978). Governance conversion makes a narrower engineering move. When experience exposes a recurring missing obligation, representation, procedure, or control, the lesson can become a model, skill, validator, permission, measurement, or other durable structure that later work inherits. The organization has then changed not only the failed instance but the environment in which the next instance will be produced.
MAGE does not assume that engineering knowledge should eventually become deterministic, or even enforced. Some judgments remain qualitative because the underlying objective is contextual, plural, or difficult to observe independently. Others contain narrower obligations that can be enforced even while the larger judgment remains open. The relevant unit is therefore the obligation, not the task.
4.4.2 Validate the Change, Not the Producer's Story
Every engineering process eventually asks the same question: what evidence is sufficient to accept this change? Code review, testing, static analysis, operational measurement, and expert judgment answered that question long before coding agents arrived. Commodity implementation does not make those techniques obsolete. It makes them more important, because changes can now arrive faster than human inspection can scale. This subsection is not a chapter on software testing. It is about how a proposed change encounters the engineered environment.
Validation begins with an obligation. If the obligation is settled, evidence can test whether the change satisfies it. If the obligation is still a hypothesis — whether users want the feature, which workflow they prefer, which tradeoff is acceptable — no test suite can manufacture the missing product knowledge. In that regime the evidence is experimental. MAGE can govern the experiment and preserve known constraints; it cannot turn an open question into a mechanical truth.
For a settled obligation, the producer's own report is rarely enough. A human says the change is correct; an agent says the tests pass; a tool announces completion. Each claim may be useful, but the strongest assurance comes from evidence generated independently of the judgment seeking admission. For consequential changes, prefer evidence that can challenge the producer's own claim rather than relying on self-report alone. Where the consequence warrants it, place admission control outside the producer's ability to redefine success.
That evidence can come from several sources, and no single kind of representation monopolizes validation. Where a model exposes a consequential property — a transition rule, a permitted edge, a declared bound — an appropriate analysis can produce evidence about whether that property holds; other models support navigation or reconstruction and need not carry an invariant at all. A model need not be an oracle to be useful. Evidence can equally come from outside a model: a test exercises a behavior, a compiler or type system rejects an illegal state, a reference implementation supplies a differential oracle, a metric reveals degradation, a human reviewer decides a property that remains judgment-laden. MAGE requires only that the obligation and the evidence be adequate to the claim.
An oracle is whatever supplies the judgment against which the observed result is evaluated. An oracle is stronger when its judgment is sufficiently independent of the implementation being judged and explicit enough that its verdict has a stable meaning. Sometimes that oracle is a simple property: the output is an ordered permutation of the input. Sometimes it is a type or schema. Sometimes it is a reference implementation. Sometimes it is a human rubric.
SOFTWARE ENGINEERINGInset — The oracle problem
Software testing has a longstanding asymmetry: executing a program can be much easier than deciding whether the result is correct. Testing research calls this the oracle problem.44. Earl T. Barr et al., “The Oracle Problem in Software Testing: A Survey,” IEEE Transactions on Software Engineering 41, no. 5 (2015): 507–25. An expected output supplies an easy oracle for some cases; complex behavior may instead require properties, reference implementations, metamorphic relations, models, runtime evidence, or human judgment.
Commodity implementation widens the asymmetry. Agents can cheaply produce implementations, variants, and test inputs, but producing more candidates does not supply an independent basis for judging them. As generation becomes cheaper, the engineering bottleneck moves toward stating what must hold and constructing evidence capable of distinguishing acceptable realizations from unacceptable ones.
Existing behavior can stand in for a surprising amount of written specification. In a compatibility port or a structure-preserving migration, a reference implementation supplies an executable oracle: challenge a new realization with the same inputs and compare it against known behavior. Tests, benchmarks, compatibility suites, and production traces close the target further. The implementation problem may stay enormous, but much less of the engineering question remains open.
Generative validation extends the oracle from a point to a space. An example test pins one case: given this input, expect that output. A property states a law over a domain and lets a generator search for a counterexample. The orientation is falsification: build the evidence to break the claim, not to confirm it. This matters because a failing example invites a local repair — patch the branch that makes that input pass and the example goes green — while a broader property keeps drawing neighbours that still exercise the law, so a repair that satisfied one input fails the next. When the defect exposes a general law, preserve the law rather than only the example that revealed it. And before treating a quiet run as evidence, ask whether the search actually reached the semantic region the claim covers: whether the generator exercised the input space's productions and drove the model's declared claims, not merely how many lines happened to execute.
The same separation appears at the organizational scale, and one industrial team converged on it independently. Docker's autonomous coding fleet builds the producer–grader separation into its architecture: the agent that writes a change never grades it. An independent reviewer with its own model invocation, a determinized test suite, and a human gate stand between every pull request and main. The fleet may open a pull request; it may not merge one.55. Manuel de la Peña, “A Virtual Agent Team at Docker: How the Coding Agent Sandboxes Team Uses a Fleet of Agents to Ship Faster,” Docker, May 1, 2026, https://www.docker.com/blog/a-virtual-agent-team-at-docker/. A practitioner writing elsewhere on the same engineering blog named what collapsing the two roles costs. He had let an agent migrate his blog and approved the result by its outcomes without reading the changes: reviewing outcomes while skipping changes "is not a security model. It's a prayer."66. Vladimir Mikhalev, “The Untrusted Autonomous Workload: How AI Coding Agents Reshape What Isolation Has to Do,” Docker, May 26, 2026, https://www.docker.com/blog/untrusted-autonomous-workload-ai-sandboxes/.
Evidence also ages. Chapter 3 gave the placement rule: evaluate an obligation at the earliest boundary where it becomes legible and enforceable, so a defect is caught close to its cause. Consequential work often deserves a second boundary: re-evaluate at the last safe point before the consequence becomes difficult to reverse. The two rules answer different questions — when can this property first be decided honestly, and what is the last point at which stale or invalid evidence can still be caught?** The security analogy is time-of-check to time-of-use (TOCTOU): a property established at one instant may no longer hold when the protected action occurs because relevant state changed in between. The problem here is broader than the classical TOCTOU race, but the engineering instinct is the same. Figure 4.4-2 draws the span between them.
The second check need not repeat every earlier check; re-establish only the evidence whose freshness matters to admission. A test result recorded hours ago describes the revision that produced it. A review verdict describes the change that was reviewed. If relevant state can change before admission, the evidence can cease to describe the artifact now crossing the boundary. Done is a claim, not a stored fact. At a consequential boundary, ensure that the evidence still justifies the consequence; where freshness cannot otherwise be established cheaply, re-derive the relevant evidence rather than trusting a stale green checkmark. The gain compounds with velocity: the more often a system ships, the more a cheap re-check at the last safe boundary is worth.
4.4.3 Operate the Environment
The governed engineering environment is itself a consequential system. Models drift. Controls become noisy. Assumptions expire. Evidence pipelines break. Gates preserve obsolete policy. Skills encode procedures that are no longer appropriate. The machinery this chapter has accumulated therefore requires the same observation, maintenance, reconciliation, and retirement it provides to the system it governs.
A method also has to survive ordinary days, not only successful changes. Repositories fill disks, branches conflict, agents disappear, queues stall, tests flake, deployments degrade, and assumptions age. Engineers often know how to recover, but that knowledge may live only in someone's head, where an autonomous agent cannot reliably reach it. MAGE therefore brings Modeling and Alignment into operation. Lifecycle models represent healthy progress and important failure states. Events make consequential transitions observable. Runbooks preserve sanctioned reactions, including where judgment remains. Metrics expose the quantities needed to judge whether the loop is healthy. Together they support an operating loop: recognize the state, surface the right reaction, take the appropriate action, and observe what happened next.
These structures play different roles. A lifecycle is a model; an event supplies evidence that a transition occurred; a runbook preserves a procedure; and a metric supplies measurement evidence. None is binding merely by existing. Constraints, validators, and gates give selected operational obligations consequence where the required state is legible and the evidence adequate. Keeping those roles separate lets the operating environment grow without collapsing representation, evidence, procedure, and enforcement into one mechanism.
A lifecycle model starts with the healthy path of a recurring activity — how a bug moves from reported to fixed, how an agent moves from dispatched to landed — then names the failure states that matter and the sanctioned recovery from each. Pair prohibitions with recovery paths. "Never use this destructive merge command" leaves a stuck agent with a prohibition and no exit; "when this state occurs, do not use X; use Y because it preserves the already-applied work" supplies the missing path. DocAble learned this when agents used a destructive Git escape during stuck merges: the durable rule named both the forbidden action and the recovery that preserved already-applied work.
Operation improves further when an important transition can make the right procedure arrive at the right moment. Do not rely on an actor to remember to inspect the lifecycle, notice that a transition happened, retrieve the right procedure, and invoke it. Where a recurring event is mechanically observable and a standard response is useful, bind the reaction to the event. A mechanical trigger does not require a mechanical response: the firing can be deterministic while the payload still contains judgment. An agent reaches a context threshold, and a hook deterministically requests a handoff the agent must still compose. A deployment fails a health check, and the event opens a recovery path an engineer must still judge. Automate the arrival of the decision procedure, not every decision.
Runbooks preserve how an experienced engineer reasons through a recurring situation, and not every step in one has the same decision structure. Type each step: Execute when the outcome is mechanically determined and can become an executable operation; Delegate when the step requires bounded judgment, supplying the relevant state, choices, criteria, and required output; Escalate when the decision remains irreducibly human, preparing the evidence and surfacing the decision explicitly. Where possible, surround a judgment-laden step with deterministic work: prepare the data mechanically before the decision, execute and measure mechanically after it, and preserve the rationale separately from the trace. That does not make the judgment deterministic; it bounds where judgment occurs while leaving the surrounding work replayable and checkable. Runbooks can drift, so their procedures need the same maintenance discipline as other engineering models.
Operational measurement should expose the state the decision actually depends on. A worker count may be a poor proxy for memory pressure; measure the pressure. A marker on disk may not establish liveness; record lifecycle state explicitly. Ousterhout's rule is apt:77. John Ousterhout, “Always Measure One Level Deeper,” Communications of the ACM 61, no. 7 (2018): 74–83. measure one level deeper — one level, not arbitrarily deep. A metric earns its maintenance cost when the deeper cut exposes an optimization target, failure class, or decision the surface number hides. A whole-project cloud bill tells you what you spent; cost by service may tell you what to change. Keep measurement separate from enforcement: a metric supplies evidence, a validator may interpret that evidence against an obligation, and a gate may act on the resulting verdict.
The operating loop will still stall. When it does, resist the reflex to add another mechanism, and first diagnose what kind of problem you have.
PRACTICEInset — When work stalls, diagnose before adding machinery
Repeated failure does not always call for another control. Before you build one, ask which part of the loop is limiting the work.
Table 4.4-1. What may be limiting The question If that is the gap Searcher Can the available intelligence find a good candidate? Change the model, tools, decomposition, or search strategy. Model Is consequential structure missing from the representation? Make it explicit. Oracle Can the environment tell good candidates from bad? Strengthen the evidence, validation, or simulation. Target Do we know what "good" even means? Stop building and learn. A stronger reasoner does not settle an unknown requirement; another validator does not supply a missing abstraction; more modeling cannot answer a product question whose answer must come from users or deployment.
Engineering capital depreciates
This chapter has now shown accumulation three times over: models entering governance, mechanisms preserving what alignment established, recovered structure joining the environment. Operation supplies the counterweight. Engineering capital depreciates. A model can stop being authoritative while remaining in the repository. A control can outlive the failure class it was built to catch. Evidence can go stale while its green marker persists. The asset that once retired a recurring cost becomes, quietly, a recurring cost itself.
So the operating rhythm includes a standing audit of the environment's own holdings. For each model, mechanism, procedure, and evidence pipeline, ask:
- Is this model still authoritative? Does it still correspond to the system it claims to describe?
- Does anything consume it? A representation nothing queries, checks, or generates from is prose with a schema.
- Does its consumer still enforce a consequential obligation? The chain from model to mechanism to consequence can break at any link.
- Is the evidence still fresh? A recorded verdict describes the past, and the ground churns under it.
- Has the assumption changed? Controls preserve the assumptions of the episode that created them.
- Is the control producing more friction than protection? Three symptoms say the return has gone negative: upkeep crowds out product work; the checks cry wolf, flagging more than they catch; or the guarded failure is both cheap and rare.
- Can the artifact be retired?
Retirement is not failure. §4.3 priced construction and upkeep together: build the smallest asset that retires the recurring cost, and retire it when the return disappears. A governed environment that only accumulates eventually governs against its own past. The discipline that converts lessons into structure must be paired with the discipline that removes structure whose lesson no longer applies.
4.4.4 Package Recurring Judgment
One kind of engineering capital remains. Models preserve recurring knowledge about the system. Controls preserve recurring obligations. Skills preserve recurring procedures and judgment. Some recurring engineering knowledge should be enforced by the environment; other knowledge properly remains available to the reasoner. That knowledge can still become engineering capital: package the representations, examples, heuristics, decision procedures, and tools that improve recurring judgment so that later work inherits them rather than reconstructing them from a generic prior.
A skill is one such package. It can range from a compact procedure resembling a disciplined prompt to a substantial body of explicit domain knowledge with structured representations, examples, correctness conditions, retrieval, and executable tools. The relevant distinction is not how elaborate the package is, but what it lets a fresh agent inherit. A useful skill gives the agent some of the abstractions, distinctions, evidence, and practices that an experienced practitioner would otherwise have to reconstruct or supply interactively. Concretely, a reusable skill needs three things: the representations needed to reason about the work, the recurring decisions that matter within it, and a procedure for making those decisions. The point is not the particular file format or vocabulary. It is to turn tacit practice into a bounded reasoning environment that a fresh agent can enter repeatedly.
Packaging does not itself enforce anything. Prose, examples, models, and decision procedures can make a satisfactory decision substantially more likely while leaving the decision probabilistic. The same skill can also invoke enforcing mechanisms or carry the models and correctness conditions those mechanisms consume. Where an obligation should not depend on the agent's reasoning, the environment still needs a constraint, validator, gate, permission, or other mechanism that can act on that obligation independently. A skill is therefore not confined to one side of the improve-judgment / enforce-obligations division: its direct instructions may support a reasoner's judgment, but it can also carry or invoke the machinery through which selected obligations are enforced.
Delivering the packaged knowledge is its own engineering problem. Loading everything all the time is the obvious solution, and often a bad one: every rule, model, example, and failure lesson competes for attention, so a larger context can make navigation harder rather than easier. Separate the small standing context that must remain continuously salient from the richer slices retrieved when the task makes them relevant, and treat the delivery as an intervention whose effect can be measured. Context delivery changes what the reasoner can see; by itself, it does not change what the environment permits. A richer slice of the right rules makes a good action easier to find; only a constraint, validator, or gate makes a bad action unavailable.
The pattern is appearing across the industry, and Chapter 5 examines several such systems side by side. Two previews suffice here. Shopify's public account describes agent sessions made visible and searchable, with repeated judgment mined into shared skills and defaults, so the environment gets incrementally smarter with each pass instead of every session starting cold. Zenseact's describes the organizational variant: a centrally owned runtime and safety envelope, with domain teams contributing agents as little more than a directory holding a config and a skill — the organization centralizes the mechanism and decentralizes the judgment.
Skills close the loop this section opened. They are another part of the distributed cognitive system that surrounds the immediate reasoner: one more place where durable engineering knowledge lives outside any single context window, alongside the models that carry facts, the evidence that records what occurred, and the controls that enforce what must hold.
4.4.5 Portable Moves
Strip away DocAble, the companion repository, and the current generation of agent harnesses, and a small set of moves remains from this chapter's dynamics and practice.
- Leave a degree of freedom open while variation is producing useful information; govern it when the choice becomes consequential and sufficiently understood.
- Establish an alignment kernel before transformations that could damage properties already known to matter.
- Use implementation as an instrument of inquiry, not as a substitute for explicit knowledge.
- In a brownfield system, recover latent structure before inventing replacement structure.
- Audit → Drain → Promote when turning an existing obligation into blocking enforcement.
- Use the weakest representation that preserves the distinction the obligation requires, and the cheapest judge adequate to decide it.
- Prefer independent evidence capable of challenging the producer's claim.
- Retain engineering structure only while the questions or failures it addresses justify its cost.
Appendix D collects these moves as bench procedures. The method itself can also become the kind of asset this section described: Appendix E shows MAGE's recurring distinctions and procedures packaged as a self-governance skill an agent can inherit rather than reconstruct.
4.4.6 Carrying Knowledge Forward
Realization amplification compressed exploration, and connection amplification compressed experience. Engineers could try more while a question remained live, and could encounter related consequences close enough together to recognize their connection. Governance conversion turned selected lessons from that experience into durable engineering structure.
Ratcheteering describes what happened next. The resulting structures remained open to evidence and correction, but the consequential concerns they governed did not silently return to implementation detail. The memory episode illustrates ratcheteering: the model changed repeatedly, but memory never returned to being somebody else's problem.
These mechanisms do not make the engineering environment progressively complete. They make it progressively capable of carrying consequential knowledge forward. Models can be revised or retired; controls can change; new evidence can reopen an old question. The gain is that consequential properties do not silently fall back from engineering knowledge into assumptions that each engineer or agent must reconstruct.
Chapter 5 changes the unit of analysis. It treats agents, models, controls, evidence, and the surrounding engineering environment as parts of a production system, and asks what kind of software factory emerges when capable realization becomes cheap.
Works Cited
- Hutchins, Edwin. Cognition in the Wild. MIT Press, 1995.
- Zhang, Jiajie, and Donald A. Norman. “Representations in Distributed Cognitive Tasks.” Cognitive Science 18, no. 1 (1994): 87–122. https://doi.org/10.1207/s15516709cog1801_3.
- Argyris, Chris, and Donald A. Schön. Organizational Learning: A Theory of Action Perspective. Addison-Wesley, 1978.
- Barr, Earl T., Mark Harman, Phil McMinn, Muzammil Shahbaz, and Shin Yoo. “The Oracle Problem in Software Testing: A Survey.” IEEE Transactions on Software Engineering 41, no. 5 (2015): 507–25.
- Peña, Manuel de la. “A Virtual Agent Team at Docker: How the Coding Agent Sandboxes Team Uses a Fleet of Agents to Ship Faster.” Docker, May 1, 2026. https://www.docker.com/blog/a-virtual-agent-team-at-docker/.
- Mikhalev, Vladimir. “The Untrusted Autonomous Workload: How AI Coding Agents Reshape What Isolation Has to Do.” Docker, May 26, 2026. https://www.docker.com/blog/untrusted-autonomous-workload-ai-sandboxes/.
- Ousterhout, John. “Always Measure One Level Deeper.” Communications of the ACM 61, no. 7 (2018): 74–83.