G.4 Expanding Delegation
Direct human review can compensate for weak modeling and alignment while knowledgeable reviewers remain available and the volume of delegated work remains manageable. A person can inspect an agent's output, reconstruct the relevant intent, notice an unexpected choice, and request a correction. That approach scales poorly when review demand grows with cheap machine realization: the cost of review continues to grow with the amount of realized work. Oversight therefore moves toward the models, evidence, validators, constraints, and gates that govern many realizations at once. Experts still review consequential work, but more of their attention can go to whether the right obligations were chosen, whether the representations remain faithful, whether the evidence is sufficient, and whether an exception is acceptable. The object of review moves; accountability does not. Part VII develops this change in sign-off as a professional consequence of agentic engineering.
This shift changes what limits scale. If every generated change requires line-by-line inspection by the person who requested it, autonomous implementation can increase output without proportionally increasing accepted engineering work. The relevant quantity is not simply how much agents produce, but how much expert attention is required to understand and govern that production. Better representations can reduce repeated reconstruction; independent evidence can settle routine questions; constraints can remove invalid choices before review; validators and gates can reject covered failures without consuming reviewer attention. The remaining review can then concentrate on obligations for which judgment still adds value. The support-apparatus ratio recorded in Part V (and recapped in Figure G.2-2) is the visible trace of this: as delegated realization grows, support apparatus can become a substantial engineering product in its own right — which is an observation about where effort accumulates, not a claim that a larger ratio is better.
This also changes where engineering effort is likely to move. When realization is expensive, substantial effort naturally accumulates in implementation. When realization becomes cheaper, the economics favor greater investment upstream in deciding what should be true and downstream in establishing whether it became true. Requirements, architecture, models, constraints, evaluation, evidence, operations, and the mechanisms that connect them become relatively more important. Implementation does not disappear, and people may still write code where doing so is useful. The organizational change is that implementation effort no longer provides a good proxy for engineering effort. Part V traces this as a redistribution of engineering effort, not the disappearance of implementation.
The boundary of safe delegation can then expand as the environment improves. A work unit that initially required an engineer to reconstruct several facts manually may become delegable once those facts are represented. A migration that once required exhaustive inspection may become easier to govern once architectural dependencies and differential evidence are available. A recurring qualitative review may acquire a structured rubric and better evidence even if it remains judgmental. Engineering capital can therefore move the delegation boundary by reducing the uncertainty or review cost attached to later work.
That boundary should not be confused with model capability. An agent may be technically capable of taking an action that the organization has no adequate basis for permitting autonomously. Conversely, a modest model can sometimes receive substantial delegated authority inside a tightly constrained environment because its possible actions are bounded and independently checked. The useful unit of adoption is consequently capability within an engineered boundary, not the advertised capability of the underlying model.
This is also why organizations should expect adoption to remain uneven. Security boundaries may justify strong enforcement early because their consequences are high and many relevant properties are mechanically evaluable. Product discovery may remain much more judgmental. A mature service may accumulate enough architecture and operational evidence for substantial autonomous maintenance while a rapidly changing prototype does not justify comparable machinery. Scientific software may govern provenance and reproducibility strongly while leaving scientific validity to domain experts. The book's theory predicts these mixed profiles rather than convergence toward one universal autonomy level, as Part VI argued.