Skip to content

§4.3 Brownfield Engineering

§4.1 closed on a convergence: greenfield and brownfield MAGE share one dynamic, discovering knowledge through realization and converting selected knowledge into durable structure. This section takes up the brownfield case directly. Its question is not whether the dynamics apply — they do — but what to do when the implementation arrived before the models: how to recover useful engineering structure from a system that already exists.

4.3.1 Start from the System You Own

Most readers do not begin on a blank page. You inherit code, tests, schemas, architectural documents, CI rules, permissions, dashboards, conventions, and operational lore. Some of that inheritance is debt; some is useful engineering structure that has become fragmented, stale, or difficult to reuse. Brownfield engineering is therefore partly an exercise in recovery: find the representations and controls the system already owns, reconnect them to the territory they describe, repair or retire what has drifted, and invest in new structure where recurring engineering costs justify it.

SOFTWARE ENGINEERING

Inset — Software evolves; its engineering structure must evolve with it

Software engineering has long treated change as a normal property of successful systems rather than an exceptional maintenance phase. Lehman's work on software evolution observed that systems embedded in a changing real-world environment must continually adapt, while their complexity tends to increase unless engineers act to control it.11. Meir M. Lehman and L. A. Belady, Program Evolution: Processes of Software Change (Academic Press, 1985).

MAGE extends that maintenance obligation beyond implementation. Models, validators, gates, runbooks, and other governance mechanisms describe or constrain a system that continues to change. They can therefore drift, preserve assumptions that no longer hold, or impose costs that no longer earn their keep. A governed engineering environment is not a finished artifact. Its engineering structure must evolve with the system it governs.

Brownfield work enters through many doors, but two are useful extremes. At one end is the fast-built system whose implementation outran its explicit engineering account: code exists, perhaps even a working product, but the architecture, obligations, and operating rules were never made durable. DocAble began near this end: a plausible product was built quickly, then migrated into a governed one. At the other end is the established organization with years of accumulated structure: ADRs, schemas, ownership files, developer portals, CI policy, tests, and operational playbooks. Its problem is not absence but fragmentation.

A fast-built system must author more of its initial models. An established system can recover and reconcile more of what it already knows. Both begin by inventorying the representations and controls already present, then investing where missing structure repeatedly consumes judgment.** A similar interpretation appears in recent developer-portal practice: a developer portal's value was never its UI but the curated context underneath — service and API catalogs, golden-path templates, ownership, CI/CD signals — and agentic engineering turns that human-facing surface into an execution substrate, "the portal isn't dying; it's becoming infrastructure".22. Red Hat, “Why Developer Portals Matter More in the Age of AI Agents,” Red Hat, 2026, https://www.redhat.com/en/blog/why-developer-portals-matter-more-age-ai-agents.

Neither starting condition requires a complete governed environment before useful work begins. Preserve the obligations already worth preserving, surface the representations already worth trusting, and let implementation and operation expose what is still missing. The method is constant; the justified starting stock is not: a brownfield system begins with the accumulated evidence of its own history, however poorly represented. Nor does the starting condition determine how settled the intent is. Greenfield does not imply uncertainty, and brownfield does not imply settled intent: a blank-page regulated system can carry highly settled obligations, while a mature product may still be exploring what its users need.

When is recovery worth the effort? Repeated failure changes the calculation. In a familiar domain a capable agent can infer adequate abstractions, architecture, and implementation from surprisingly little explicit structure. When agents keep reconstructing the same concept, cross the same hidden boundary, or re-search the same unproductive region, another attempt is worth less than changing how the problem is represented. Implementation-first becomes churn when the search repeatedly reconstructs structure whose explicit representation would cost less than continued rediscovery. In a brownfield system, that structure is often already present — expressed again and again in the implementation, but never named.

4.3.2 Recover Latent Models

Begin with two questions, because representation and enforcement are different things. First, what does the organization already reason through? Which documents, schemas, graphs, tests, or configurations repeatedly save someone from reconstructing the system from raw implementation? Second, what is already enforced? Which permissions, types, validators, CI checks, reviews, or operational gates currently constrain or admit work, and what can each actually decide? A mature system may already contain substantial local Alignment even when its system-level representations are weak. A schema may reject invalid data, a type system may exclude states, and CI may block a merge. The brownfield problem is therefore not simply to build a better map and add controls afterward. Strengthen representation and enforcement together: connect useful representations to the territory they describe, make important obligations explicit enough to evaluate, and place enforcement where the required semantics are available.

Brownfield models need not begin with an engineer inventing the right abstraction in advance. Existing implementation often contains repeated structures that point toward models the system has never named. Static analysis, search, clustering, and other detectors can surface those patterns at scale, but their findings are evidence rather than conclusions. Repetition may reveal a missing domain concept; it may instead reflect appropriate low-level implementation, a protocol boundary, test scaffolding, or ordinary local cleanup. The detector finds places worth examining. Engineering judgment determines whether the repeated structure carries a shared meaning worth representing.

DocAble encountered this problem through a primitive-density detector, which identifies interfaces expressed heavily through bare strings, integers, tuples, and other low-level values. Primitive density is intentionally a weak signal: such interfaces sometimes indicate that a concept is being repeatedly reconstructed rather than represented explicitly, but they are also common in perfectly legitimate code. The detector therefore produced hundreds of findings without claiming that hundreds of models were missing. Most findings needed ordinary annotations or local cleanup, or occurred in test fixtures and protocol adapters where primitive-heavy interfaces were appropriate.

A smaller subset behaved differently. Across related document mutators, the same parameter shapes recurred: language mutations repeatedly carried the same language and provenance concepts; alt-text mutations repeatedly carried description and attribution state. Looking across those instances revealed something that no individual finding established on its own. The repeated primitives were not merely similar code. They were recurring expressions of domain concepts that the implementation had never named.

Repetition alone is not enough to make that inference. The finding must be interpreted in its architectural context. A primitive-heavy external protocol adapter is different from primitive-heavy internal domain code: the former may be faithfully representing an external interface, while the latter may be forcing every consumer to reconstruct a concept the system already understands. The same distinction applies more generally. A recurring structure may be legitimate shape, mechanical debt, or evidence of a latent concept. The useful question is therefore not simply does this pattern repeat? but what explains the repetition, and would naming the shared meaning change future reasoning?

This gives a practical route from weak implementation signals to explicit models. Surface repeated structure; partition findings by architectural role and likely cause; compare related instances for shared meaning; name a concept only when that meaning earns a representation; then migrate consumers and enforce whatever narrower relationships the new model makes decidable. Figure 4.3-1 summarizes that process. The schematic is intentionally simpler than the judgment that produces it: detection narrows the search space, while engineering interpretation decides whether there is a model to recover.

Discovering models during a brownfield migration: weak signals locate work, partitioning separates it, and only repeated shape that earns a name becomes a model A top-down flow. An existing codebase is surveyed by weak discovery signals — primitive density, repeated shapes, and a component or boundary view. Partitioning the findings splits them three ways: legitimate primitive-heavy shape is accepted or exempted; mechanical debt is fixed by codemod; and a latent concept, a repeated semantic shape, earns a name. Naming the model produces a typed vocabulary and explicit relations, which drive consumer migration and then a narrower enforcement of types, lints, and a drift guard. The result is a stronger model substrate on which the next pass runs. Weak signals locate the work; they do not dictate the remedy. EXISTING CODEBASE the territory weak discovery signals primitive density repeated shapes component / boundary view PARTITION FINDINGS LEGITIMATE SHAPE protocol · tool · boundary · test MECHANICAL DEBT annotation · cleanup LATENT CONCEPT repeated semantic shape accept / exempt leave as-is codemod mechanical fix NAME THE MODEL the concept earns a name typed vocabulary JobId · ChunkId · MutationSpec … explicit relations ownership · boundary · behavior … MIGRATE CONSUMERS codemods + agents NARROWER ENFORCEMENT types · lints · drift guard STRONGER MODEL SUBSTRATE next pass Weak signals locate the work; they do not dictate the remedy. Only repeated structure that earns a name becomes a model — the rest is accepted or mechanically cleaned.
Figure 4.3-1. Discovering models in a brownfield system. Weak signals — primitive density, repeated shapes, the component and boundary view — identify places worth inspecting; they do not dictate the remedy. Partitioning separates legitimate primitive-heavy code and mechanical debt from repeated shapes that reveal missing engineering concepts. Naming those concepts as typed vocabulary and explicit relations changes the substrate on which later recovery and enforcement operate.

Field record — Primitive-density recovery

The primitive-density backlog fell from 698 file-level findings on June 8 to 472 on June 9 and 299 on June 10. The declining count is not the important result. What matters is that the remaining population was repeatedly reclassified as it shrank:

  • June 8 — 698 findings, partitioned into cluster-uniform, mechanical, per-site-judgment, boundary, and test-fixture-or-exempt buckets.
  • June 9 — 472 findings.
  • June 10 — 299 findings, dispositioned as annotate, exempt, extract-a-model, redesign, retire, or defer-for-judgment.

Some findings disappeared through mechanical work, some exposed missing shared types, some warranted exemptions or lint changes, and some required design judgment. The chronology records those observations; it makes no claim that a single planned wave caused each step of the decline.

4.3.3 Reconcile the Map and the Territory

A recovered representation is a hypothesis. The inherited architecture document, the induced concept, the reconstructed dependency structure — each claims to describe the system, and each may be stale, wrong, or asking its question at the wrong grain. Until a recovered model has been reconciled against the implementation, it should not govern anything.

Begin with audit. Join the representation to the implementation and surface disagreement without blocking work. A legacy map is expected to drift; making every inherited claim binding before reconciling it would merely enforce old errors. Repair disagreements through ordinary feature and maintenance work without assuming in advance that prose or code is correct: the representation may be stale, the implementation may be wrong, or the representation may be asking the question at the wrong grain. Promote selected obligations only after the representation is trustworthy enough for the consumer that will rely on it. Add structure when retrieval, analysis, generation, validation, measurement, or gating will use it; do not enrich a model merely because more structure can be represented. The companion From-Drifted-Wiki-to-Trusted-Model operator card holds the drill whole.

To make the shape concrete, suppose an inherited wiki describes the system but has drifted from the code, so agents and engineers repeatedly re-derive what is actually true from source. Audit the map against the implementation and reconcile disagreements through ordinary feature and maintenance work: sometimes the prose is stale; sometimes the code is wrong. Promote only claims trustworthy enough for the consumer that will rely on them. The result is a queryable model later work can use, while unresolved claims remain advisory rather than acquiring unearned enforcement.

A disagreement has the familiar three-way diagnosis — the representation is wrong, the grain is wrong, or the implementation is wrong. Do not decide which in advance; shifting a legacy codebase from "the tests pass" to "these properties hold across all behaviors" surfaces disagreements by design.

SOFTWARE ENGINEERING

Inset — When the model and the code disagree

Tighten a model against real code and it will, sooner or later, contradict the code. The disagreement usually points to one of three questions:

  • Is the model wrong? It may claim something the code never promised. Refine the model.
  • Is the model at the wrong grain? It may be asking a question at a level of abstraction that cannot see the answer. Build a different model at the grain the property lives at.
  • Is the implementation wrong? The model may be right, the code disagrees, and the disagreement may expose an implementation defect.

The discipline is to decide which before you act. Fixing code when the model was wrong bakes in a false claim; refining the model when the code had a bug hides the bug behind an adjusted spec.

Once the diagnosis is made, the vocabulary of §4.1 applies unchanged: active alignment brings realization and model into correspondence, and the reconciled model then remains in the environment as passive alignment, preserving what reconciliation established.

A valid future invariant is not always a valid present admission rule. A legacy system may violate a worthwhile new rule everywhere on the day you introduce it. Making the rule blocking immediately does not improve the system; it simply stops existing work. Introduce the obligation in three beats: Audit, Drain, Promote.

The move — Audit → Drain → Promote

  • Audit. Install the proposed rule as evidence only. It reports every detected violation and blocks no commit, so it never breaks an in-flight agent.
  • Drain. A fix wave repairs the existing findings under sufficient tests and other evidence to preserve intended behavior, converting the legacy code to the new convention behavior-for-behavior.
  • Promote. Once the governed surface satisfies the obligation, enforce it at the boundary so detected violations cannot re-enter.

Cloudflare's public account describes a similar approved → enforced sequence, landing each policy control advisory and flipping it to enforced once the tree is clean, so the promotion never breaks work in flight.33. Timo Reimann, “How Cloudflare Enforces Engineering Standards Using Ai,” Cloudflare, August 4, 2026, https://blog.cloudflare.com/engineering-standards-enforcement/. Figure 4.3-2 draws the three beats as one lint is worn into a legacy tree.

Squash, zero, promote: how a lint is worn into a legacy tree one step at a time A new smell-detecting lint moves through three states, left to right. First it lands audit-only: it reports every finding but blocks no commit, so it never breaks an in-flight agent. Then a fix wave — an agent making behavior-preserving edits under test coverage — drains the finding count to zero; this is a loop that runs under the audit-only state until nothing is left. Only when the count reaches zero is the lint promoted to blocking: it now fails any commit that reintroduces the smell. A one-way barrier sits after the blocking state, marking that the class of smell can never return. Squash the findings, drive them to zero, then promote — and the smells that leave stay gone. new lint Audit-only reports, blocks nothing Findings → 0 a fix wave drains it Blocking fails a regressing commit land at zero behavior-preserving fixes, under test coverage a detected violation now blocks re-entry
Figure 4.3-2. Audit, Drain, Promote. A new lint lands audit-only — every finding reported, no commit blocked, so it never breaks an in-flight agent. A fix wave drains it to zero, and only then is it promoted to blocking, so future violations within the mechanism's declared detection surface are refused at that boundary.

The final step requires care. DocAble did not promote primitive density itself into a universal blocking rule. The detector was a smell: useful for locating work, but too broad to define correctness. Once the findings had been understood, narrower mechanisms could prevent the specific regressions that mattered. Sometimes you promote the obligation, not the instrument that discovered it. An audit may reveal an obligation that deserves enforcement while not itself being the right mechanism to provide it; the detector can remain advisory while a narrower type, constraint, validator, or gate enforces the obligation.

4.3.4 Invest Where It Pays

Brownfield MAGE is not a project to model everything. The inherited system is large, the recoverable structure is larger than the structure worth recovering, and every model, validator, and gate carried forward must be maintained against a system that keeps changing. Prioritize the places where explicit structure retires a recurring cost:

  • Knowledge is repeatedly reconstructed. Engineers and agents keep re-deriving the same facts from raw implementation.
  • Recurring failures expose missing structure. The same class of defect returns because nothing represents or enforces the obligation it violates.
  • Consequential properties are hard to evaluate. A property that matters cannot be checked without a representation at the right grain.
  • Fragmentation blocks useful joins. The facts exist, but they live in disconnected representations that cannot answer questions together.
  • Recurring judgment is expensive enough to convert. A repeated human decision has become stable and frequent enough to encode.

Even then, a recurring gap does not automatically deserve machinery. Price the expected future loss against the cost of building and carrying the response. Frequency and consequence are useful first approximations; detectability, reversibility, legal exposure, and the stability of the obligation can move the decision substantially. Cheap, rare failures often remain judgment. Cheap, frequent failures may justify a lint, script, or reusable procedure, because repeated agent and human attention is itself cost. Consequential failures justify a larger investment and a stronger assurance response: better representation, prevention by construction where possible, stronger sensing and validation, and enforcement at an appropriate boundary. Catastrophic or legally intolerable failures need not recur even once to earn that investment.

Price the upkeep as well as the construction. Models drift, validators require calibration, and skills age. Engineering capital depreciates: a control that once retired meaningful cost becomes overhead when the obligation disappears or the upkeep exceeds the return. Build the smallest asset that retires the recurring cost; retire it when the return disappears. Fleet-scale brownfield work amplifies the same economics: Spotify reports a migration workflow in which work that previously involved hundreds of teams over weeks could be scoped by one engineer over days, with tooling targeting and scheduling concurrent agent work.44. Max Charas and Marc Bruggmann, “Honk: Autonomous Code Migration at Spotify,” Spotify Engineering, 2026, https://engineering.atspotify.com/.

Recovering one model is useful: a concept the system had been expressing for years gains a name, a representation, and perhaps an enforced obligation. Repeating the process changes something larger — the environment in which subsequent engineering occurs.

Works Cited

  1. Lehman, Meir M., and L. A. Belady. Program Evolution: Processes of Software Change. Academic Press, 1985.
  2. Red Hat. “Why Developer Portals Matter More in the Age of AI Agents.” Red Hat, 2026. https://www.redhat.com/en/blog/why-developer-portals-matter-more-age-ai-agents.
  3. Reimann, Timo. “How Cloudflare Enforces Engineering Standards Using Ai.” Cloudflare, August 4, 2026. https://blog.cloudflare.com/engineering-standards-enforcement/.
  4. Charas, Max, and Marc Bruggmann. “Honk: Autonomous Code Migration at Spotify.” Spotify Engineering, 2026. https://engineering.atspotify.com/.
© James C. Davis, 2026–present