Alignment
Premise. Guidance shapes behavior; enforcement determines what the environment accepts.
Alignment connects the epistemology of Validation to the control problem of Delegation. Validation asked what evidence justifies engineering action. Delegation asked what consequences a capable but fallible agent should be authorized to produce. Alignment connects these questions: when an engineering obligation matters, how should evidence about that obligation govern what the engineered environment permits or accepts?
Software engineers already build this kind of control into their environments. A linter checks some properties while code is being written. A pre-commit hook can reject a change before it becomes a commit; a pre-push hook, before the work leaves the developer's machine. Continuous integration can evaluate the assembled change before merge. A deployment gate can inspect a release candidate before production accepts it. Runtime controls can restrict what the deployed system is permitted to do. These mechanisms operate at different boundaries because they can see different things: a linter may decide a property from one source file, an integration test may require the assembled system, and some properties become visible only after the system is running.
There is also a difference between asking for a property and checking it. A coding standard can tell an engineer not to introduce a forbidden dependency; an architectural check can reject that dependency. The first influences the producer. The second gives engineering knowledge consequences.
MAGE calls this broader practice Alignment: connecting engineering obligations to mechanisms that can constrain work, produce evidence, evaluate that evidence, and control what the environment accepts.
Alignment Principle. Make engineering obligations enforceable by encoding them into mechanisms that constrain actions, produce evidence, evaluate that evidence, and control admission.
Check it where it can be decided¶
Consider the familiar progression: edit → commit → push → merge → deploy → runtime. Different checks attach to different points. Why not run every check as early as possible? Because some properties do not yet exist in a form that can be decided.
Suppose a deployed web application must correctly render data emitted by a backend service. No source-file check can establish that property: the source may compile, unit tests may pass, the emitted data may be individually valid, and the question becomes answerable only when the pieces compose in the served system. The same problem appears within a single task: when a change must end with the code and its architectural model in agreement, disagreement during intermediate commits may be perfectly legitimate. Rejecting each intermediate commit would enforce the right obligation at the wrong boundary.
This gives us a placement rule: check an obligation at the earliest boundary where you can actually decide it. Earlier checks usually make failures cheaper to repair, but moving a check earlier than its evidence permits does not make the control stronger. It makes the check incapable of deciding the property it claims to enforce. The boundary follows the property.
Guidance is not enforcement¶
Suppose an agent receives the instruction: do not merge unless the security tests pass. Placed in a prompt, project instructions, or a skill, this guidance may strongly influence behavior. But the reasoner can misunderstand it, forget it, or decide incorrectly that the condition has been satisfied. Connect the same condition to a merge gate and something different happens: the environment determines whether the merge proceeds.
Both are useful, and Alignment is not an argument against instructions, documentation, review, or expert judgment. It asks which engineering obligations matter enough, and are sufficiently evaluable, that satisfying them should not depend solely on each future producer remembering and interpreting them correctly.
An obligation the work can be wrong against¶
Before mechanizing an obligation, distinguish three questions that software engineering often mixes together:
- Correspondence — do two representations agree? Does an architectural model match the dependency graph extracted from the implementation?
- Conformance — does an artifact satisfy an independent obligation? Does this PDF satisfy the relevant standard?
- Acceptance — will the receiving environment take the result? Will CI admit the change?
Evidence for one does not establish the others, and agreement is not correctness. A representation derived from code may perfectly describe an unauthenticated endpoint: the representation and the implementation agree, and the security obligation is still violated. Alignment therefore needs an obligation against which the work can actually be wrong — a requirement, invariant, policy, tolerance, permission, schema, or model. What matters is not whether the representation was hand-written or derived; it is whether the mechanism has an independent engineering condition against which to judge the work.
Four roles, and a preference for prevention¶
Once the obligation and its boundary are known, we can ask what the environment should do. Four roles are useful:
- Constraint — narrows what may happen.
- Sensor — observes what happened and produces evidence.
- Validator — evaluates evidence against an obligation.
- Gate — controls whether work may cross a boundary.
These are roles, not four separate tools or four sequential stages. A CI test can sense behavior, validate the result, and gate the build; a narrow interface can constrain available actions while also producing evidence about their use.
The first engineering preference is prevention: if an invalid state can be excluded cheaply and reliably, prefer excluding it to repeatedly detecting it later. A closed enumeration can make an illegal value unrepresentable; a permission can prevent an agent from invoking a consequential action at all. Not every property can be held structurally — some appear only across sequences of individually legal actions, or only in the assembled artifact or at runtime. Those need evidence: observe, validate against the obligation, and decide whether the verdict controls admission. The mechanism follows the property — a type, an architectural check, a model checker, human review, and a deployment test are not stronger and weaker forms of Alignment; each answers a different engineering question.
From blacklists to sanctioned paths¶
Generative implementation makes one Alignment problem visible: an agent can invent implementation paths its designers did not anticipate. Suppose all document color mutations must derive from a canonical color model. Banning every known bypass works while the forbidden surface stays small and recognizable, but future implementations may combine otherwise legitimate operations into a bypass no blacklist anticipated. Sometimes the allowed path is easier to characterize: the sanctioned abstraction attaches provenance ordinary callers cannot create, and the mutation boundary requires that provenance before accepting the operation. The rule of thumb: when forbidden paths are enumerable, ban them; when the allowed path is easier to characterize than all possible bypasses, make admission depend on evidence of the allowed path.
Not everything should be enforced¶
Alignment is not a march toward making every engineering decision mechanical. An engineering concern can stop in several places:
- Residual — the concern is not represented adequately; an engineer or agent must reconstruct the relevant meaning.
- Judgment required — the obligation is explicit, but no available evaluator can decide it adequately.
- Evaluated only — evidence is produced and evaluated, but the result does not control admission.
- Governed — the environment constrains the action or makes admission depend on the verdict.
These are design choices, not maturity levels. A cost metric may be worth observing without a hard budget. A probabilistic validator may serve triage while remaining too uncertain to block production. Stronger Alignment does not mean more gates; it means the obligations engineering chooses to enforce are enforced dependably at appropriate boundaries.
When failures become controls¶
Alignment also explains how an engineering environment changes over time. A failure may initially require diagnosis and judgment: an engineer discovers that a consequential obligation was absent, weakly represented, checked at the wrong boundary, or left to guidance when the environment could have enforced it.
When the same judgment is likely to matter again, the engineer can change the environment. A recurring review question becomes a validator. A convention becomes an architectural constraint. A remembered check becomes a gate. A known-dangerous operation disappears behind a sanctioned interface. MAGE calls this governance conversion: converting engineering knowledge acquired through experience into durable control.
The result has value beyond the individual failure that produced it. Future engineers and agents no longer need to reconstruct the same judgment from scratch, and the environment can prevent or reject whole classes of recurrence. These accumulated models, constraints, validators, gates, and sanctioned paths are a form of engineering capital: prior engineering judgment embedded in reusable structure.
But conversion is not automatic. Some failures expose obligations that remain difficult to represent or evaluate, and some judgments should remain judgments. The question is not "Can we add another gate?" It is whether a recurring engineering judgment can be represented faithfully enough, evaluated reliably enough, and placed at an appropriate boundary to deserve authority over future work.
From engineering knowledge to engineering control¶
The three units now fit together. Agents asked how engineers delegate realization without delegating responsibility: BOUND → EQUIP → AUTHORIZE → VERIFY. Modeling asked how engineers preserve consequential distinctions while leaving irrelevant choices free, and how those purposeful reductions remain interpretable, connected, and correspondent to the system. Alignment asks what happens when some of that engineering knowledge must do more than inform the next reasoner: state the obligation, check it where it can actually be decided, choose an appropriate mechanism, and determine what happens when the check fails.
The environment does not remain fixed. Failures reveal what it does not yet know, see, evaluate, or enforce, and recurring judgments can sometimes be converted into durable control. The next unit, Failure-Aware Engineering, asks how to make that learning systematic.
The objective is not to eliminate engineering judgment. It is to decide where judgment should remain judgment, and where recurring judgment should become durable engineering structure. That is Alignment.
Materials¶
- Lecture 1 slides — From Guidance to Authority (forthcoming) — (coming soon)
- Lecture 2 slides — Governing Realization (forthcoming) — (coming soon)
Readings¶
The alignment principle
- MAGE §3.1, "Where Obligations Can Be Enforced." and MAGE §3.4, "Growing the Governed Environment." Davis, 2026. The chapter opening and §§3.1–3.3.3 establish the core Alignment argument: guidance versus enforcement, correspondence / conformance / acceptance, the earliest decidable boundary, the four control roles, and matching mechanisms to properties. §3.4 then shows how the governed environment changes through experience: failures expose missing controls, recurring judgments can undergo governance conversion, and accumulated controls become engineering capital. Read these sections before class so that the session can use the vocabulary to reason about concrete engineering situations rather than spending the session introducing it. Full citation: James C. Davis, Model-Based Agentic Engineering, 1st ed. (2026), https://davisjam.github.io/model-based-agentic-software-engineering/book/mage-book/index.html, chap. 3 introduction, §§3.1–3.3.3, and §3.4.
Place the check where the semantics exist
- Saltzer, Reed, and Clark, "End-to-End Arguments in System Design" (1984). Chapter 3 acknowledges the connection explicitly. This classic systems argument — place a function where the semantics needed to decide it exist, not merely as early or as low as possible — is the placement rule in an older register. Read it asking which of the paper's communication-system examples transfer to engineering checks, and what plays the role of the "ends" when the system is an engineering environment rather than a network. Full citation: Jerome H. Saltzer et al., “End-to-End Arguments in System Design,” ACM Transactions on Computer Systems 2, no. 4 (1984): 277–88, https://doi.org/10.1145/357401.357402.
Control as an engineering tradition
- Works by Nancy Leveson, TODO.
The systems-safety and control tradition establishes that constraints plus evidence, and feedback plus intervention, are serious pre-agent engineering ideas rather than vocabulary invented for AI. The specific work and portion to assign are still being selected.
The organizational precedent
- Simons, "Control in an Age of Empowerment" (1995). The organizational precedent. Management faced the delegation problem long before software agents existed: how to grant people real autonomy while keeping the organization's consequential obligations intact. Simons' answer is a designed system of controls, not more instruction — empowerment and control engineered together, the same pairing this unit makes of capability and authority. The reading also guards against a misreading: Alignment is not a distinctively AI-era invention. Full citation: Robert Simons, “Control in an Age of Empowerment,” Harvard Business Review 73, no. 2 (1995): 80–88, https://hbr.org/1995/03/control-in-an-age-of-empowerment.
Optional / further reading
- MAGE, Chapter 3, §§3.3.4–3.3.7 and MAGE §3.5, "When Controls Become a System.". The later §3.3 sections elaborate mechanisms that land better after the walk from linter to CI to constraints to sanctioned paths — provenance-carried admission in particular is taught in class rather than required beforehand. §3.5 answers what happens after a hundred controls accumulate: the natural extension for an interested student, not required preparation. Full citation: James C. Davis, Model-Based Agentic Engineering, 1st ed. (2026), https://davisjam.github.io/model-based-agentic-software-engineering/book/mage-book/index.html, §§3.3.4–3.3.7 and §3.5.