§2.5 System Knowledge: Connecting the Models
The preceding sections drew different models of the same product because they asked different questions about it. Together, the models expose structures through which engineers can reason about DocAble: its components and dependencies, permitted behavior, ownership, policy, quantitative bounds, and realized history. The system keeps them as a family of structured representations, each built for an engineering purpose, connected by shared identity and traceability where engineers repeatedly need to relate one view to another. It would be natural to collapse these views into one large model, but DocAble has no such model, and MAGE requires none.
2.5.1 A Heterogeneous Modeling Substrate
DocAble's models differ in shape on purpose. Some are registries of components and their relations. Some represent state transitions. Some encode contracts or decisions. Others capture quantitative relationships or realized history. Roughly 130 structured model records plus several data dialects represent component identity and metadata, deployment topology, inter-service edges, data flow, required environment, state machines, and the governance graph itself. They share identities, derivations, query surfaces, and traceability where those relationships are useful, but nothing forces them into a single representation. The models are represented as data rather than importing the implementation they describe, keeping model and territory separable enough for correspondence to be checked. When a model is meant to provide independent evidence of correspondence, keep it distinguishable from the artifact it checks. Chapter 5 returns to this modeling substrate as part of the governed engineering environment and shows how it is maintained and used in practice.
The unity that matters is semantic, not syntactic. A service named in a structural registry is the same service a flow model references. A work item in an ownership representation is the same work item whose lifecycle and measurements appear elsewhere. A provenance event identifies the operation whose effect it records. Nothing demands that these models share a file format or metamodel; they need only agree on the identities required to connect them (Figure 2.5-1).
That agreement follows one discipline: prefer shared identity and derivation over independently maintained copies of the same fact. Where one model needs a fact owned elsewhere, join through a stable identity or derive a projection. For example, the declared edge model is held in exact correspondence with the deployment edge set, and deployment then derives its access policy — one invocation grant per declared edge — from those edges rather than maintaining a second endpoint list; a journey model joins to those endpoints by call site rather than copying them. Where duplication is unavoidable, make the correspondence explicit and checkable. This reduces the number of independent truths that must be reconciled by hand.
Static and runtime representations need not collapse into one model either. A computation model can identify remediation work and its declared relations while a provenance record identifies the mutations realized during one run. Shared identity allows observations to join the structural model when a question requires it: DocAble associates staged-run latency and cost with computation identities. The per-session edit record remains separate because it answers a different question — what happened to this artifact — and is a separate provenance reduction rather than a per-run execution graph over computation nodes.
What Happened, and What Evidence Records It
That realized history deserves its own reduction. A production system may emit millions of events — logs, traces, metrics, audit records — while still making a consequential question difficult to answer: what changed this artifact, which operation changed it, and what evidence supports that account? A provenance model selects and structures the history needed to answer such questions. For DocAble, the consequential history is the history of changes to the user's document, so it records consequential mutations as structured operations: each operation names its target, the mutation made, the remediation pass responsible, and the evidence retained about the change (Figure 2.5-2). The record is narrower than a general execution trace and more directly useful for explanation, debugging, and audit.
Choosing the grain is part of the modeling decision. "Document modified" is too coarse to explain a consequential change; a complete execution trace leaves later reasoning to reconstruct which events matter. DocAble instead records consequential mutations as typed operations and derives audit, debugging, and changelog views from that history.
During a PDF remediation session, these mutations are recorded as typed edit values in a per-session log, and a replay engine can apply the same sequence deterministically. The representation can therefore explain what happened and reproduce the resulting mutations while omitting the reasoning, analyses, control decisions, and retries that produced them.
The computation graph asks what computations may compose. The provenance record says which consequential mutations occurred in this run. Replay asks whether that realized history can be reproduced (Figure 2.5-3).
One limit is worth stating. A provenance model does not guarantee that every consequential operation actually leaves the required record. DocAble addresses that separately, by requiring typed mutation verbs to emit attribution and checking that obligation structurally. That check is a hard presence guarantee over a soft content claim: it can establish that every mutator leaves an attribution record, but not that the semantic account inside that record is correct — that remains the author's. This chapter models the evidence an operation should leave; Chapter 3 asks what mechanism enforces that obligation.
Questions That Span Models
A model makes selected properties of a system available for reasoning. Joining models makes properties that cross those representations available for reasoning. The structural model can answer a structural question, the deployment model a deployment question, the decision model a permission question — but some engineering obligations span them.
DocAble's sharpest worked case is the interservice call edge. A typed edge records that one service may call another — but the same edge is joined to deployment configuration to determine the environment variable carrying the callee's URL, to the runtime identity a deployment executes under to derive the invocation grant the cloud's access-control configuration must contain, and to the service catalog to check that the documented topology agrees. Twenty-eight modeled call edges produce ten distinct invocation grants once shared runtime identities are accounted for. No single source artifact contains that deployment plan; it is derived by joining models. DocAble modeled the join after living the corresponding failure: service URLs were once wired correctly while the matching invocation grant was absent, and staging went down with authorization errors. The repair did not add another bespoke grant. It made URL wiring and authorization two consequences of the same modeled edge.
Connections can create information that no individual model contains. The call model does not contain the environment-specific runtime identity; the deployment environment does not independently state the intended call topology. Joining them determines the permission that should exist.
The same pattern reaches from models to their evidence. Each invariant in DocAble's composed state-machine model declares its facts: the participant lanes that touch a coordination point, the coordination primitive, and the temporal shape of the property. A derivation over those facts — never a hand-typed label — assigns the invariant a verification tier; the tier determines what kind of checker the invariant must cite; and a check joins each citation against the checker corpus on disk, failing when the mandated checker is missing. A declared property can therefore determine what kind of evidence must exist. The join makes the obligation derivable. Chapter 3 discusses the separate decision of whether such an obligation should block a commit.
These crossings are also where information is lost. Each representation can be locally reasonable while their composition fails to preserve a distinction the engineering decision requires. The entitlement failure of the decision section is one example: the system possessed the fact that an entitlement was unlimited, but the representation presented to the admission decision had reduced that richer state to a scalar. A join through shared identity does not remove the risk; it gives the crossing a name and a place where the correspondence can be checked.
Shared identity also provides a query entry point. An agent — or an engineer — enters the substrate through the object its task names and traverses outward, following identities across models rather than loading everything. A canonical query surface lets the agent traverse represented relationships rather than rediscover them through repository search: name a component and ask what it owns, what it depends on, which flow reaches it, what its lifecycle looks like. The agent does not need the whole model. It needs the connected slice that answers the engineering question in front of it. Figure 2.5-4 draws that entry and outward walk.
2.5.2 Maintain Explicit Correspondence
Multiple models create a correspondence problem: represented facts can diverge from the system or from one another. Engineer the relationship between map and territory so divergence is caught rather than assumed away. Which machinery you reach for depends on which side is the source of truth.
Make the direction of correspondence explicit. Where the implementation is the source of truth, derive the model. Where the model is the source of truth, generate the downstream artifact. Where neither fully determines the other, maintain traceability and check the mechanically decidable correspondences. Figure 2.5-5 lays the three cases side by side.
Derive when implementation owns the truth
The model is a projection of the code, reconciled at build time. A component-zone registry reads the real directory tree; a journey's dependency list is induced from its real call sites. The drift check is a reconciler — it re-derives the model from the code and fails on divergence. This pattern is common when models are introduced into an existing codebase. ### Generate when the model owns the truth
Code, configuration, or documentation are emitted from the model. The service catalog and the wire-contract types are generated from the service-flow model. The drift check is a freshness and provenance check — the generated artifact carries a header, and a hand-edit or a stale regeneration is a finding.
DocAble's remediation subsystem illustrates both directions around one source of truth. Its explicit computation model names remediation work and its composition, and the execution machinery consumes that model to determine dependency order and readiness. The model therefore governs composition operationally: this is the generate case in the consequential sense that realization follows the model rather than reconstructing composition from node implementations.
A separate Python remediation-graph view projects that structure for analysis and governance. For Office formats, its edges derive from the same registry read/write metadata used by scheduling; for PDF, parallel typed declarations are held in correspondence by blocking checks. The view can therefore be queried, joined to other models, and checked for drift without itself becoming the runtime scheduler.
One source of truth can thus support both directions at once: execution consumes the authoritative model, while analytical representations derive from it or from parity-bound declarations around it. The important question is which representation owns the fact and which consumers are required to follow it.
The graph is derived, but it remains a view: it suppresses method bodies, internal algorithms, runtime history, timing, and most document state, keeping only the entities and relations its structural questions need. Derivation prevents one kind of drift by construction — an engineer never authors the edges.
Trace and check where neither owns the truth
Where neither side can be fully derived from the other, the move is traceability plus drift checking. Relate each model element to the implementation that realizes it, then check the correspondences a machine can decide. Keep that relation live: resolve the target against current code when the check runs, rather than storing a line number or a symbol name that is itself another snapshot able to drift. A deleted target then surfaces as a finding instead of a stale green edge. In DocAble a traceability gate resolves each model-to-code anchor through a language-aware resolver — one for each source language — and reports every anchor that no longer resolves. Before the gate existed, such renames could silently strand model references.
Some correspondence remains semantic
Mechanical correspondence has a declared surface. Some claims are decidable: an anchor resolves, an observed dependency belongs to the permitted set, or a generated artifact is current. Other claims remain semantic and require judgment; unmodeled properties have no model-based coverage. A recurring miss is a candidate for refining the representation or adding a control (Figure 2.5-6).
For invariant-bearing models, one useful executable pattern is a structured schema plus checked predicates over that schema. The schema defines entities, relations, fields, or states; the predicates state properties that a checker evaluates over the declared domain. Other executable models may instead drive generation, simulation, transformation, or analysis. A schema alone does not establish correspondence to a changing system; checks or derivations are needed where that correspondence matters. Figure 2.5-7 draws that relationship.
None of this proves the model right. Synchronization does not establish that the model is correct, that the architecture it expresses is good, or that the modeled surface captures every relationship that matters. A pointer can resolve while its meaning has changed; a schema and its generated client can agree while a live producer emits something else. The guarantee the machinery offers is narrower, and worth stating plainly:
For the surface the model claims, check the correspondence that can actually be decided.
That boundary returns in Chapter 3, where correspondence and correctness part company. Derive where implementation is the source of truth; generate where the model is the source of truth; trace and check where neither fully determines the other.
2.5.3 Modeling Makes Properties Available to Engineering
A model does not govern the system merely by existing. A structural model can name a forbidden dependency without preventing it; a behavioral model can name an illegal transition without rejecting it; a measurement model can represent a budget without stopping work that exceeds it. Modeling makes these properties explicit. It does not enforce them.
DocAble's quantitative model makes the point concretely. Usage and pricing determine an attributable GenAI cost, read against a declared budget; concurrency determines a capacity demand, read against a capacity envelope. The wording of the property is deliberate: the model says cost can be compared against a budget, not that cost must not exceed one. DocAble attaches no hard invariant to either quantity — cost is routed to the administrative pane for observation and accounting, while capacity is controlled through a runtime-tunable limit. In the configuration examined here, these quantities are observations rather than enforced bounds. A cost figure is evidence; whether crossing a budget should block work, raise an alarm, or simply be recorded is a separate question of enforcement.
Figure 2.5-8 draws that separation: the model defines the quantity and the reference; whether a comparison is merely observed, used to adapt behavior, or enforced by a gate is Alignment's decision. The model that makes spending attributable does not decide the budget, guarantee that the numbers are collected correctly, or enforce the bound.
Local syntactic obligations can often be checked with little explicit modeling. Architectural boundaries, temporal obligations, distributed ownership, provenance requirements, and end-to-end policies require enough representation to expose the relevant system relationships. Modeling makes those relationships available to analysis and, where appropriate, to governance.
They are also available to design. A dependency graph may reveal that an execution boundary is artificial; a behavioral model may expose a recovery protocol that should change; a quantitative model may show that an otherwise plausible decomposition cannot meet its resource envelope. In each case, analysis of the model can change the architecture before implementation commits to the choice.
Alignment does not require every obligation to have a separately represented model. Some properties are already observable at an action or artifact boundary and can be constrained directly. Models extend the range of properties available for analysis and enforcement.
DocAble's models arose around engineering questions that implementation alone made expensive or unsafe to reconstruct: where mutation may occur, how remediation computations compose, who owns in-flight work, which services may communicate, whether a user is entitled to an operation, what bounds apply, and what happened during a run. Different questions required different reductions; shared identities connect them where a decision crosses views.
Chapter 3 asks which obligations the engineered environment should enforce, where they become decidable, and what evidence and mechanisms are sufficient to enforce them.
AIDEEP DIVEBeneath commodity intelligence: why structure helps
MAGE deliberately treats intelligence as a commodity. The method should not depend on a particular neural architecture, training recipe, or context-window size. Still, the language-model systems available in 2026 offer a useful view one level below that abstraction. They suggest why engineering the structure around an intelligent agent can matter as much as improving the intelligence itself.
Access to state is not the same as reasoning over state. Modern language models can accept large contexts, but their ability to use those contexts depends on the reasoning the task requires. OOLONG, for example, was designed to require semantic processing and aggregation across much of a long input rather than retrieval of a few relevant passages. Frontier models degrade substantially on these tasks even when the input fits within their nominal context windows.11. Amanda Bertsch et al., “Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities,” 2025, https://arxiv.org/abs/2511.02817. Recursive Language Models (RLMs) make the distinction especially clear. Instead of placing a long input directly into the model's context, an RLM keeps it in an external environment that the model can inspect programmatically, decompose, and pass in pieces to recursive model calls. On several long-context tasks, this restructuring substantially improves performance, including on inputs far beyond the underlying model's context window.22. Alex L. Zhang et al., “Recursive Language Models,” 2025, https://arxiv.org/abs/2512.24601.
This belongs to a broader shift from treating a language model as a function that produces an answer toward treating it as a reasoner operating within an environment. ReAct interleaves reasoning with actions that obtain new observations from an external environment;33. Shunyu Yao et al., “React: Synergizing Reasoning and Acting in Language Models,” in “International Conference on Learning Representations,” special issue, International Conference on Learning Representations, 2023. CodeAct extends this idea by allowing executable code to serve as an expressive action language.44. Xingyao Wang et al., “Executable Code Actions Elicit Better LLM Agents,” in “Proceedings of the 41st International Conference on Machine Learning,” special issue, Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, vol. 235 (2024): 50208–32. RLMs go further by making context management and decomposition themselves part of the interaction with the environment. The common lesson is not that any particular harness is universally correct. It is that the structure through which intelligence encounters a problem changes what that intelligence can effectively do.
Recent work offers a possible explanation in terms of compositional generalization. Zhang, Li, and Khattab hypothesize that frontier models may already be highly capable on many bounded problems while remaining poorly "managed" on long-horizon ones: a larger task succeeds when it can be decomposed into smaller tasks that individual model calls can reliably solve.55. Alex Zhang et al., “The Mismanaged Geniuses Hypothesis,” 2026. Zhang and Khattab subsequently argue that a harness can supply some of this structure by transforming complicated external state into observations that remain locally familiar to the underlying model; their RLM experiments suggest that such structure can improve generalization across both task length and domains.66. Alex L. Zhang and Omar Khattab, “Language Model Harnesses Are Compositional Generalizers,” 2026. These are developing results rather than a settled theory of language models, but they offer a useful substrate-level interpretation of an engineering observation: more intelligence and more context do not remove the need for structure.
MAGE operates one level above this work. A generic agent harness can manage context, call tools, delegate subtasks, and decide how to decompose computation. Software engineering also provides semantic decompositions of the system itself. An architecture model exposes components and boundaries. A data-flow model exposes consequential paths. An automaton exposes legal state transitions. Ownership and obligation models expose ownership and constraints. Bidirectional links among models and implementation let an agent move from code to the relevant model, across related models, and back to the affected code rather than reconstructing all of that structure from a flat collection of source and prose.
This distinction also marks the boundary between Modeling and Alignment. Modeling gives commodity intelligence representations in which consequential engineering relationships are explicit and traversable. Alignment makes selected obligations enforceable by connecting them to mechanisms in the engineered environment—hooks, validators, tests, gates, runtime checks, and others—that can enforce them. Better learned decomposition may improve how agents perform work inside that environment. It does not decide which customer, organizational, safety, or societal obligations the environment should preserve.
The connection is therefore explanatory, not foundational. Future models may reason over much longer contexts, learn better decompositions, and require different harnesses. MAGE should survive those changes. Its engineering claim is more durable: capable reasoning becomes more useful when consequential system structure is made explicit, and consequential obligations become more reliable when the environment enforces them. Most importantly, those structures keep engineers in control of what the system is for, what it must preserve, and which tradeoffs are acceptable, even as increasingly capable machines take on more of the work required to realize those decisions.
Models are useful without enforcement. They preserve decisions, expose relationships, support analysis, and let engineers and agents reason from shared representations rather than reconstructing the same knowledge from prose or code. Alignment adds another use: some modeled properties can become obligations that constrain what the engineering process may admit. We turn to that next. Chapter 5 returns to these models as part of the engineering environment surrounding agentic production, and Chapter 6 examines when carrying explicit representations is worth their cost.
Works Cited
- Bertsch, Amanda, Adithya Pratapa, Teruko Mitamura, Graham Neubig, and Matthew R. Gormley. “Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities.” 2025. https://arxiv.org/abs/2511.02817.
- Zhang, Alex L., Tim Kraska, and Omar Khattab. “Recursive Language Models.” 2025. https://arxiv.org/abs/2512.24601.
- Yao, Shunyu, Jeffrey Zhao, Dian Yu, et al. “React: Synergizing Reasoning and Acting in Language Models.” In “International Conference on Learning Representations.” Special issue, International Conference on Learning Representations, 2023.
- Wang, Xingyao, Yangyi Chen, Lifan Yuan, et al. “Executable Code Actions Elicit Better LLM Agents.” In “Proceedings of the 41st International Conference on Machine Learning.” Special issue, Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, vol. 235 (2024): 50208–32.
- Alex Zhang et al., “The Mismanaged Geniuses Hypothesis,” 2026.
- Alex L. Zhang and Omar Khattab, “Language Model Harnesses Are Compositional Generalizers,” 2026.