Skip to content

§5.2 Inside DocAble's Software Factory

This section illustrates

✓ The Modeling Principle · ✓ The Alignment Principle · ✓ Governance Conversion · ✓ Residual Human Judgment

DocAble was built over roughly 20 weeks by one engineer (me) directing a fleet of coding agents. By August, a routine day involved six to eight agents working in parallel and roughly 200 commits landing in the repository. The resulting system included 491,090 lines of production code, 1,501,907 lines of tests and other support apparatus, 35,323 lines of infrastructure code, and 28,507 lines of explicit system models.

New work did not originate primarily in the coding agents or the factory's existing models. It came from engineering needs: my own goals for the system, user feedback, failures observed in production, changes and obligations in accessibility standards, public guidance from accessibility advocacy groups,11. National Federation of the Blind, “Creating Nonvisually Accessible Documents,” 2014, https://nfb.org/images/nfb/products_technology/creating_accessible_documents.docx. and the demands introduced as the system moved into institutional use. I decided which of these pressures became engineering work and made the consequential decisions about requirements, architecture, validation, and deployment.

The existing engineering environment then shaped how that work could be realized. Architecture, tests, system models, project instructions, validators, deployment machinery, and accumulated controls carried decisions made during earlier work. Agents performed most of the implementation inside that environment. I inspected almost none of the code they produced. My attention went instead to deciding what should change, resolving consequential design questions, examining evidence and failures, deciding what work could be admitted, and changing the production system when later work should inherit an earlier judgment.

That division of work is the factory examined in this section. The section moves in seven steps: the problem the factory was built for; the build, read from the repository's own weekly record; the mature factory's representations and machinery; that factory in operation; the supervisor's changed job; the economics; and the liabilities that came with it. The chronology matters because none of it arrived fully formed. As realization accelerated, failures exposed missing representations, controls, and supervisory surfaces; the production system changed in response.

The requirements were not unusual for production software. Chapter 1 stated them as R1–R8, and this section repeats them where they begin to do their work. What was unusual was the implementation capacity available to one engineer.

5.2.1 The Problem

In 2024, the U.S. Department of Justice required state and local governments, including public universities and schools, to make covered web content and mobile applications conform to WCAG 2.1 Level AA.22. U.S. Department of Justice, “Nondiscrimination on the Basis of Disability; Accessibility of Web Information and Services of State and Local Government Entities,” 2024. The original compliance date for larger public entities was April 24, 2026. Four days before that deadline, the Department extended it by one year.33. U.S. Department of Justice, “Extension of Compliance Dates for Nondiscrimination on the Basis of Disability; Accessibility of Web Information and Services of State and Local Government Entities,” Federal Register 91 (2026): 20902–12, https://www.federalregister.gov/documents/2026/04/20/2026-07663/extension-of-compliance-dates-for-nondiscrimination-on-the-basis-of-disability-accessibility-of-web.

The Department's explanation described a production problem. Public entities held large amounts of inaccessible material and limited staff to remediate it. Correspondence cited in the rule argued that current technology, "including generative AI," could not reliably automate the remediation of STEM materials at scale. The Department itself concluded that the advancement and availability of technology "did not meet the Department's expectations."

The scale was easy to see. Consider one ordinary 16-week university course producing three slide decks and two PDF readings each week. At three hours to remediate a deck and five hours to remediate a PDF, one course's materials cost

16(3×3+2×5)=304 hours≈7.6 work weeks

The five-hour figure is conservative — one colleague reported spending a week on a single PDF. A university offering 3,000 courses would face roughly 912,000 hours under the same assumptions. The affected educational system is much larger than one university: NCES counted 99,297 operating public elementary and secondary schools in 2023–24,44. National Center for Education Statistics, “Table 3. Number of Operating Public Elementary and Secondary Schools, By School Type, Charter, And State or Jurisdiction: School Year 2023–24,” 2025, https://nces.ed.gov/ccd/tables/202324_summary_3.asp. and 1,570 public degree-granting institutions in 2022–23.55. National Center for Education Statistics, “Table 317.20. Degree-Granting Postsecondary Institutions, By Control and Classification of Institution and State or Jurisdiction: Academic Year 2022–23,” 2024, https://nces.ed.gov/programs/digest/d23/tables/dt23_317.20.asp.

The Department also recorded a quality concern: institutions facing this workload might respond with rapid procedural box-checking rather than sustainable accessibility work. Passing an accessibility checker is not the same as making a document accessible — or producing work an institution could defend in a lawsuit.** Digital-accessibility litigation is tracked continuously rather than reported as a closed figure. UsableNet's ADA Accessibility Lawsuit Tracker enumerates federal and key state-court digital-accessibility cases on a monthly cadence, and recorded 4,928 web-accessibility lawsuits filed during 2025.66. UsableNet, “ADA Accessibility Lawsuit Tracker,” 2026, https://info.usablenet.com/ada-website-compliance-lawsuit-tracker. Laura Carlson maintains the higher-education-specific tally of lawsuits, complaints, and settlements.77. Laura Carlson, “Higher Ed Accessibility Lawsuits, Complaints, And Settlements,” 2026, https://www.d.umn.edu/~lcarlson/atteam/lawsuits.html. Both are living sources; the count above is the reading at the access date recorded in the bibliography, not a settled total.

I repeat the eight requirements from §1.2 here deliberately (Table 5.2-1). There they defined the problem DocAble set out to solve. Here they explain why demonstrating semantic reasoning was only the beginning.

Table 5.2-1. The eight requirements again, as §1.2 stated them — the same R1–R8, now read as the specification a factory had to satisfy.
RequirementDocAble must…
R1 · Multi-format remediationRemediate educational documents in the formats people actually use, including Office (PowerPoint, Word, Excel) and PDF.
R2 · Semantic remediationPerform semantic repairs that require understanding document content, not only syntactic, deterministic transformations.
R3 · Preserve the documentRetain content, appearance, and behavior except where remediation requires a change.
R4 · Standards-grounded accessibilityEvaluate and repair documents against their governing obligations (WCAG 2.1 AA), not proxy checker scores.
R5 · EvidenceRecord what accessibility requirement motivated each change and what evidence supports the resulting document state.
R6 · ReviewabilityMake outputs auditably traceable: consequential automated changes are inspectable and, where appropriate, reversible.
R7 · Production operationProcess real documents at practical sizes, costs, and latency.
R8 · Institutional deploymentOperate as a deployed service that scales to thousands of faculty at one institution, and preferably nationwide.

These requirements explain why the resulting system grew large. Semantic reasoning was only one requirement of eight. A frontier model could supply a capability that had previously been missing; it could not preserve documents, produce reviewable changes and evidence, establish conformance with the governing standard, or operate a dependable institutional service. Those obligations generated the engineering system around the model.

The first question was whether frontier models could supply the missing capability.

5.2.2 Building the Factory

The Feasibility Probe

DocAble began with a feasibility question. Document remediation requires semantic reasoning. A document is not merely a container of text: figures, tables, reading order, hierarchy, emphasis, and visual arrangement all carry meaning. Making that document accessible requires recovering enough of that meaning to present it through another channel — to a screen reader, to a keyboard, or to a student who cannot see the figure a lecturer drew.

Automated accessibility had historically been unable to supply much of this reasoning. Documents come in too many formats, and humans express the same ideas in too many ways, for conventional automation to enumerate the possibilities. Earlier reasoning systems could not reliably bridge that semantic gap. Modern frontier models were obvious candidates.

I gave a multimodal model pictures of mathematical equations, technical slides, and handwritten notes. It had no trouble interpreting them and producing useful transcriptions and alternative text.

Generative AI supplied the missing semantic reasoning. The rest was engineering.

That distinction became measurable later in the build. One of PDF remediation's central reasoning problems is recovering a document's semantic role sequence in reading order. A frontier model working on that problem from scratch cost $4.41 to $8.12 per document, and its agreement with ground truth ranged from 0.564 to 1.000 depending on which document it was given. The same class of model, given DocAble's engineered intermediate representation as a scaffold, scored 0.936 for $0.11 on the one document measured under both conditions — cheaper and more accurate than the $4.41 from-scratch run, which scored 0.818 on that same document. Cheaper models reading the same representation cost $0.06 to $0.65 while retaining between 55 and 90 percent of the frontier's measured quality.

Four qualifications matter. The experiment measured structure inference, not complete document remediation. The scores are semantic similarity of the recovered role sequence, and they say nothing about where elements land on the page; on that geometric axis the scaffolded model reached no comparable figure. The ground truth was authored by a model of the same family, so the top of the from-scratch range is self-agreement rather than an independent oracle. The two cost figures also do not share an accounting basis: the from-scratch charges are metered, the scaffolded ones a documented token estimate. The direction survives that mismatch, since under the most generous assumptions available the scaffolded path remains three to thirty-three times cheaper, but the ratio itself is approximate.

The result also varied across document types. Richer representations helped prose, and could over-merge some formula-heavy or scanned material, so the shipped system varies the representation by document regime. The experiment does not rank models. Frontier capability was already sufficient to perform difficult semantic reasoning. Engineering determined how much of that reasoning had to be purchased, what structure the model could inherit instead of rediscovering, and where different representations were appropriate.

Cost was only one reason not to make the model the remediation system. DocAble instead separated semantic judgment from consequential mutation. Models reasoned over representations; typed operations changed the artifact; independent validators checked properties that could be established mechanically; and provenance was embedded with consequential mutations, so that the eventual change record could be reconstructed from the document rather than trusted as a model-generated account.

The Two Rules of Agentic Software Engineering

I began with two rules.

Reading the code is the last resort. Not reading the code was deliberate. If building DocAble required me to inspect the implementation produced by coding agents, then implementation throughput would remain bounded by my capacity to read code. The point was not to become faster at code review. It was to build an engineering environment in which I could direct work, understand the system, evaluate evidence, and decide whether changes were acceptable without routinely reconstructing those answers from implementation.

This did not mean that I refused to read source code. It meant that needing to read it pointed to a missing supervisory surface: a model, test, lint, instrument, architectural boundary, or other representation of the property I actually needed to reason about. Reading the code was the escape hatch, not the operating model.

The resulting system should be good, fast, and cheap. This was partly a requirement for DocAble and partly a test of agentic software engineering itself. I wanted to know whether agents could help build a production system that was not merely fast to develop, but good enough to trust, fast enough to matter, and cheap enough to deploy at scale. I therefore made all three requirements of the factory.

A fast factory producing poor software would fail the test. So would a high-quality system whose development or operation remained prohibitively expensive. Agentic implementation was valuable only if the engineering system around it could achieve all three.

The Build, in Stages

I opened an empty repository and began with PowerPoint. The first program walked the contents of a slide deck and sent figures and equations to a model for description. Colleagues asked whether it could handle PDFs, so I generalized the system to Word and Excel and began adding PDF support (R1); documents from other people immediately exposed assumptions my own files had not. Putting the system online — DocAble runs in beta at scholaccess.com — introduced authentication, deployment, institutional use, and feedback from documents I had never seen (R8). Users then found outputs that passed existing accessibility checkers while losing semantic equivalence, so I replaced those checkers as the primary validation mechanism with validation grounded directly in the accessibility standards (R4). When the Department of Justice extended the compliance deadline, the immediate deadline pressure receded, and I shifted the project toward production hardening (R7). Most new work entered the factory through pressures outside the fabricator: new users and documents, observed failures, deployment demands, standards obligations, and my own engineering objectives. I translated those pressures into engineering work; the accumulated production system shaped how agents could realize it. Figure 5.2-1 gives the sequence. In a few months, a feasibility experiment had become a production software system.

The seven build stages as a dated timeline of pressure and consequence A retrospective vertical timeline of seven build stages, each row naming the stage, its dates, the pressure that forced it, and the major engineering consequence that followed. Feasibility, March: vision-language capability became usable, so the build began. PowerPoint, March to April: colleagues supplied their own decks, producing the first real remediation pipeline. Format expansion, April: Word, Excel, and PDF broke one-size-fits-all assumptions, so the architecture differentiated. SaaS, April: real users on real documents forced authentication, quota, cost, security, and exposed fidelity limits. Standards, April to May: the need for a defensible accessibility claim forced standards-grounded validation. Hardening, May to June: time became available for structural work, producing models, tests, gates, and provenance. Deployment, iteration, and optimization, July to September: operating the production system exposed performance, cost, reliability, and architectural limits, forcing repeated measurement, redesign, optimization, and re-platforming. The stages are labels drawn afterward, not a preplanned roadmap. STAGE THE PRESSURE THE CONSEQUENCE 1 Feasibility March a probe in a meeting Vision-language capability became usable → the build began 2 PowerPoint March–April the first format Colleagues supplied their own decks → the first real remediation pipeline 3 Format expansion April Word · Excel · PDF New formats broke the one-size-fits-all assumptions → the architecture differentiated 4 SaaS April a deployed service Real users on real documents → auth, quota, cost, security — and the first fidelity limits 5 Standards April–May a defensible claim The need for a defensible accessibility claim → standards-grounded validation 6 Hardening May–June structural work Time became available for structural work → models, tests, gates, provenance 7 Deployment, iteration, and optimization July–September operating what shipped Production operation exposed limits on performance, cost, reliability, and architecture → repeated measurement, redesign, optimization, and re-platforming Stages are labels drawn afterward, not a plan followed at the time.
Figure 5.2-1. The Seven Build Stages. Retrospective sequence from the March feasibility probe to the July–September period of deployment, iteration, and optimization, giving the dates, the pressure, and the major engineering consequence at each stage. The serverless re-platforming is one episode inside that final period rather than a stage of its own.

My role was concentrated elsewhere. I decomposed work, made architectural and design decisions, examined failures, decided which failures were local and which exposed a weakness in the system or its engineering environment, and changed that environment when the same judgment should not have to be made again.

The Build, Measured Weekly

The stage narrative above is retrospective. The repository's own record supplies a check. A history-mining pass reconstructed weekly series across the project's 29 weeks — commits landing on main, work units opened and closed, control artifacts, and source volume by category — with the book's earlier snapshot counts retained as validation points inside the series. Four traces build the argument: how much was realized each week; what work flowed through realization; what controls accumulated beside it; and how large the tests and supporting infrastructure became relative to the product.

The first trace is realization itself: commits landing on main per week (Figure 5.2-2). The trace contains three regimes and one recording change. The prototype barely registers — a few dozen commits in the opening weeks, two weeks with none at all. Week 5 is the ignition: 828 commits, and the series never again returns to prototype scale. Throughput then climbs through mechanization to a peak of 3,332 commits in week 12 and holds four-digit weeks through most of the mature era. Around week 22 the composition shifts — agent-attributed commits fall and automation commits rise — because a merge-train squash became the landing mechanism and each squash records many underlying agent commits as one. The shift is a change in how work is recorded, not a drop in activity: the total series is the cross-regime-comparable one, and even squash-collapsed, mature weeks land 800–1,300 commits. One caveat: nearly every commit is authored under the human git identity, so provenance is reconstructed from commit trailers and subject prefixes, not the author field.

2026-09-25T11:43:35.953179 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
Figure 5.2-2. Realization Over Time. Weekly commits landing on main across the project's 29 weeks, split by reconstructed authorship provenance, with dotted build-stage boundaries. The late-era shift from agent-attributed commits to merge-train squashes records a change in landing mechanism, not a slowdown: one squash stands for many underlying agent commits, so the total series is the cross-regime-comparable trace.

The second trace is independent of commits and immune to the landing-mechanism shift: it measures Epics, the factory's unit of formalized engineering intent (Figure 5.2-3). An Epic was not an Agile story point, a single agent invocation, or a commit. It was a committed engineering record for a substantial desired change: a product requirement, a production failure, an architectural change, a new control or model, or an improvement to the factory itself. Its template required the rationale and the cost of not doing the work; identification of the system models and invariants the change would study, update, or add; a cost and risk estimate; and a phased implementation plan with explicit acceptance criteria. Before inventing a solution, its Phase-1 design also had to examine the established approaches, libraries, tools, and schemas appropriate to the problem, and record the alternatives considered. An Epic could close only after its implementation, documentation, models, lints, tests, and other controls had been brought into agreement and independently reviewed.

Phase 1 was always design. The resulting design document encoded the proposed structure and its invariants before implementation began, so that the design could drive both the work and its eventual validation rather than documenting decisions after the fact. For nontrivial work, an independent model reviewed that design before implementation. The remaining phases decomposed the design into dispatchable work, potentially across several agents operating in parallel; an independent Definition-of-Done review closed the other end of the process.

One point in Figure 5.2-3 can therefore stand for a considerable amount of realization, though not uniformly so. The structure the template permits is one Epic to several phases, one phase to several agent dispatches, and one dispatch to several commits. Whether a given Epic used all of that structure is a separate question, and the repository answers it. Joining closed Epics to the commits whose messages cite them — 444 of 478, a join that needs no appeal to timing — puts the median closed Epic at two commits, with seven at the seventy-fifth percentile and a mean of six that campaign Epics past a hundred commits pull upward. The median Epic changed about 344 lines across roughly fifteen files and closed within a day, and across the same population the median change to production source was ten lines. Most of what a typical Epic changed was its own design document, phase notes, models, and controls; product code was the minority.

Two things keep that median from contradicting the multiplicity above. The count is a count of landings rather than of work: from week 22 a merge train squashed each agent's work space into a single commit on main, the recording change the first trace already showed, so a late Epic's two commits can stand for two dispatches rather than two units of work. The count is also a lower bound in a second sense, because the join requires a commit message to cite its Epic by qualified path, which 93% of closed Epics do; a looser join that accepts a bare name puts the median at four. Read honestly, most Epics were small, the tail was large, and the unit measures formalized intent rather than implementation volume.

An example from the larger end of that distribution shows what the multiplicity buys. One September Epic began when a user's lecture deck exposed an accessibility checker that incorrectly flagged some diagram and text-bearing shapes as lacking alternative text. Resolving it involved root-cause analysis, an independent design review, three implementation dispatches, and a final Definition-of-Done review. That is at least six agent runs behind the three landed commits its implementation phase records, and behind one opened and eventually one closed point in the weekly trace. Other Epics changed Purdue beta billing policy, made the agent registry authoritative for tracking work in flight, hardened the automated cherry-pick machinery, and added an operator-facing analyzer for production failure hotspots.

From week 12, when the per-Epic file convention was introduced, the factory created and closed Epics at sustained volume: roughly 50 to 115 opened and 11 to 48 closed per week, for seventeen consecutive full weeks. That first week records 111 openings, which is a formalization event rather than a week of extraordinary ideation: every Epic already tracked informally acquired its file at once. Throughout, "opened" and "closed" are repository events rather than workflow-status transitions: the week an Epic's file first landed in version control, and the week it first landed in the closed section of the tree. The two series do not run at the same level: 1,368 units were opened against 479 recorded closed, creation outpacing closure by roughly three to one. Work kept arriving — experiments produced information, users and unfamiliar documents produced requirements, failures exposed missing controls and representations, the architecture became visible enough to improve — and I decided, week by week, which of it to pursue. A queue being drained closes what it opens; this one did not. Part of the gap is the census window, since units opened in the closing weeks had no time to close, and part is work superseded or absorbed rather than formally closed; the record does not separate the two, so read the ratio as a shape and not as a completion rate. No conventional backlog existed to drain: DocAble was greenfield and exploratory, and much of the work that would turn out to matter had not been discovered yet, let alone listed. An Epic records work at the moment engineering judgment made it concrete enough to pursue, so the series measures the cadence at which engineering intent was formalized into named units. The factory did not run out of engineering intent. As realization accelerated, building the system continually revealed more things worth doing. That reading sharpens the division of labor this section keeps returning to: the supply of consequential engineering direction stayed a human responsibility, and abundant realization executed it once specified. This is the concrete texture behind "one engineer directing a fleet" — intent formalized at a rate no one person could implement.

2026-09-25T11:43:36.470228 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
Figure 5.2-3. Work Over Time. Epics — DocAble's unit of formalized engineering intent — created and closed per week from the convention's introduction in week 12. A record of intent formalized, not of a backlog drained: creation outpaces closure by roughly three to one. The first week's 111 openings are a formalization event, not a burst of ideation; the trace is invariant to the late-era landing-mechanism change.

The Controls Accumulate

The third trace measures one family of controls that accumulated beside that throughput: project-specific lints and gate scripts. An earlier draft of this section counted them at four repository snapshots; the weekly series (Figure 5.2-4) shows what snapshots cannot — the shape of the accumulation. The lint-file count is flat at zero through week 8, takes off sharply across weeks 9–11, and then grows near-monotonically to roughly 990 files at the mining head. The growth is cumulative: it continues regardless of that week's realization volume (the weekly correlation between commits and lint files is approximately zero), which is what distinguishes deliberate investment from a side effect of throughput. The lint-file definition is validated against this chapter's earlier counts — 595 files at the hardening-window revision, matching exactly; the companion gate-script and system-model traces use the mining's own documented definitions and are independent re-derivations rather than reproductions. Count alone says little about quality. The more useful repository evidence is provenance: 208 commits carry a paired fix-and-lint tag, marking a repair that landed together with a check against its recurrence, and 27 lints name the dated incident that motivated them.

2026-09-25T11:43:36.969922 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
Figure 5.2-4. The Controls Accumulate. Project-specific lint files (definition validated against this chapter's snapshot counts), gate/check scripts (independent definition), and system-model files at each week's last commit. The four-snapshot growth curve of the earlier draft, resolved into a 29-week series: flat through week 8, a sharp takeoff at weeks 9–11, then cumulative growth decoupled from weekly realization volume.

The Support Ratio

The fourth trace weighs the tests and supporting infrastructure against the product they support. Divide the source counted as support — tests, modeling, orchestration, and governance tooling — by production source, week by week (Figure 5.2-5). The weekly series resolves the dynamics the earlier dated snapshots compressed: the ratio begins below parity in the prototype weeks, crosses parity around weeks 5–6, climbs steeply through mechanization, and settles on a plateau between roughly 3.5 and 3.9 across the mature era. Call this the support ratio. Two honesty notes qualify the figure. First, the weekly series is an independent re-derivation with its own category boundaries, not the chapter's original line-count configuration — the two earlier snapshot ratios the mining pass could recover, one post-prototype and one at the canonical August snapshot, are overlaid as diamonds so the definitional gap stays visible; the trend (below parity, crosses, settles around three-to-one) is the comparable claim, the levels are the mining's own. Second, the stacked composition shows where the support mass lives: excluding documentation, tests dominate (about 1.75 million lines at the final week), followed by infrastructure, governance controls, orchestration, and system models. The ratio is descriptive, not a target, and support code is not capital merely because it exists — it earns that name only while future work inherits useful capacity from it. The figure shows what accumulated, not its current return.

2026-09-25T11:43:37.496008 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
Figure 5.2-5. Factory Capital and the Support Ratio, Weekly. Source-tree lines of code by category (left) and the support-to-production ratio (right) at each week's last commit — an independent re-derivation with its own category boundaries, with the two recoverable earlier snapshot ratios overlaid as diamonds. The trajectory matters more than any single value: below parity in the prototype weeks, parity crossed around weeks 5–6, a plateau near three-to-one in the mature era.

Read in order, the four exhibits form one sequence. Realization became abundant early and stayed abundant. Named engineering intent kept flowing through that capacity at a sustained cadence. Controls accumulated beside the throughput as a deliberate investment, decoupled from any week's realization volume. And the supporting software — tests, models, controls, orchestration — grew until it outweighed, several times over, the production source it existed to govern. One longitudinal case cannot establish that this sequence is necessary; it can establish that in this case it happened, in this order, and the record makes it unmistakable. What that accumulation was — and what it subsequently made possible — is the business of the rest of the section.

What Accumulated, and Why

What had accumulated? Not only implementation, but a factory around it. Some of it was ordinary software engineering. Tests accumulated. Static analyses checked properties that mattered. Deployment gates became stronger. Architectural boundaries became explicit. Other parts existed specifically because agents were doing the implementation: task briefs, agent-facing project instructions, role-specific dispatch, mechanisms for selecting the context relevant to a task, and controls over concurrent work.

Model selection was itself an engineering decision: frontier models were reserved for judgments where their additional capability earned their additional cost — the cheapest-adequate-judge rule of §4.4 applied to the fabricator rather than to the validator.

The resulting arrangement made parallel implementation cheap while admission to the production system remained serialized through evidence. Realization capacity could fan out without granting each fabricator independent authority to admit its work.

The distinction was not clean, because the two sides interacted. A model of the software architecture could guide an agent before it worked and support a deterministic check afterward. A failure discovered by an agent could expose an architectural weakness whose repair also made future agent work easier to control. The engineering environment and the product developed together.

And none of it was present when I started. In April, a production failure led me to introduce an explicit state model for a processing lifecycle. In May, duplicated descriptions of the system were consolidated into a shared modeling substrate. By July, properties represented in those models were determining what verification the repository required. Later failures led to additional models and controls for service interactions, deployment, entitlements, and resource use.

The sequence matters because I did not begin with MAGE, or with a plan for an agentic software factory. I began with a software problem and capable implementation agents. The engineering structures described in this chapter appeared as I encountered the limits of directing those agents without them — which is what the control curve's takeoff at weeks 9–11 records from the outside.

What these exhibits measure is what the eight requirements cost. The model solved the part of the problem that had historically resisted automation. It did not solve the software engineering problem around it.

Field note — The counterfactual

One counterfactual rests on direct engineering judgment rather than repository data. Within the actual constraints of this project—one primary engineer, an academic budget, and the development interval reported here—DocAble was not feasible for me without generative AI. That is not a claim that a conventional engineering organization could not have built the system. The relevant counterfactual is the project that actually existed: under those constraints, I would not have attempted a production system at this scope without commodity implementation capacity.

A second counterfactual concerns sustained delegation: GenAI made DocAble feasible; the machinery that became MAGE made delegation at this rate feasible. Once autonomous change volume exceeded anything I could inspect conventionally, raw model capability was no longer enough. Continuing at that rate required increasingly explicit representations of the system, independent evidence, constraints, validators, gates, and operating machinery through which the work could be understood and governed.

This is a within-case engineering judgment, not a controlled causal estimate.

DocAble did not solve abundant implementation by finding a way for its engineer to read implementation faster. It changed what the engineer had to supervise. What outgrew the product is where the change lives: representations that carry consequential knowledge, controls that give selected obligations consequences, and a recurring conversion of failures and repeated judgment into both. The next two subsections examine the factory those investments produced — first what it knows and enforces, then how it runs — and the subsection after them returns to the supervisor, because the accumulation's deepest effect was on that job.

5.2.3 Inside the Factory

Source code alone stopped being an adequate engineering interface as DocAble grew. Repository search could locate a symbol; it could not say which services participated in a workflow, which component owned a piece of state, or what had to stay bounded as a path changed. This subsection walks the interior of the factory the weekly series just measured: the representations the mature factory maintains, the joins that connect them, and the one representation that took over execution.

What the Factory Knows

At the repository revision examined here, DocAble contained 130 declared system models — the count the factory's canonical census query (repo-query models kinds) returns, and the census Chapter 2's "roughly 130" summarizes — each a typed, queryable representation registered in one place, most carrying their own consistency checks. The exact inventory changes with the system: the census reported here was re-derived from the live repository while this chapter was being written, and it had already drifted by one from the count taken at an earlier revision. That drift is the point. These are maintained engineering artifacts that move with the system they describe, not documentation frozen for a presentation.

The population is heterogeneous because the engineering questions are. By declared form, the 130 comprise 50 registries, 17 instruments, 13 budgets, 11 schemas, 10 constraint sets, 8 dataflows, 7 topologies, 4 state machines, 4 decision tables, 4 policies, and 2 interaction models. By target, 78 model the product, 51 model the agent fleet and its engineering environment, and one models the execution environment itself — the factory represents both the thing it produces and the system that produces it. Figure 5.2-6 draws the census; the mining pass that produced it reproduces the published count, the per-form split, and the per-target split exactly at the examined revision.

2026-09-29T20:19:01.490018 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
Figure 5.2-6. The Model Census. The 130 declared system models at the examined revision, by declared form (left) and by target (right). The factory represents both the product it produces (78) and the system that produces it (51, plus one model of the execution environment). A snapshot of a live census: regenerate at press time rather than treating 130 as frozen.

The point of the census is what those models afford. Chapters 2–4 argued that a model is a purposeful reduction matched to an engineering question. What the earlier chapters could not yet show is the column a factory adds: the supervisory consequence — what the engineer stops having to reconstruct from implementation because the representation carries it. Table 5.2-2 shows six representative specimens from the live census, chosen for the variety of question, form, and consumer.

Table 5.2-2. Six specimens from the model census. Each row: the engineering question, the representation that answers it, what consumes it, and what the supervisor no longer reconstructs from implementation.
Engineering questionRepresentationConsumerSupervisory consequence
Can a job always terminate legally?Composed cross-service state-machine model with per-invariant factsExhaustive-state checkers; the workers' own transition tablesLegal progress and termination are reasoned over states and transitions, not a multi-thousand-line worker
How much memory can a job need?Peak-memory sizing model that upper-bounds a job and derives the deployment memory tierDeployment configuration; regression tests on the boundA resource bound is inspected directly rather than reconstructed from allocation behavior
Which services exist, and what may call what?Deployment topology plus a typed inter-service call-edge modelDeploy machinery; agents localizing changesOwnership and boundaries are read from the map, not rediscovered per change
May this user submit this document?Entitlement decision table — the single policy authority for the admission decisionThe front door at upload timeThe admission rule is one table; changing policy is editing a model, not hunting scattered conditionals
What may touch a document format's internals?Registry of external seams plus typed format models owning all mutationBuild-time ban checks on raw library accessA format-corruption class is excluded by construction rather than hunted after release
Which step changed this document element?Per-mutation provenance stamps projected into a derived changelogThe customer-facing changelog; forensic queriesDiagnosis is a lookup over evidence, not a cross-file investigation

Walk any row and the same pattern appears. The lifecycle model lets its supervisor reason about legal progress and termination rather than reading the worker that implements them; the memory budget lets the supervisor reason about a bound rather than reconstructing allocation behavior; the entitlement table preserves an admission-policy distinction as one authoritative representation rather than logic diffused through handlers. Two million lines of source sit under this factory. The recurring refrain of this section is that the supervisor is not supervising two million lines through two million lines — supervision runs through representations sized to the questions that matter.

The Joins: What the Factory Can Derive and Enforce

A repository can contain excellent documentation and still supervise nothing. What separates DocAble's models from documentation is that the environment joins them and acts on them: representations connect to other representations, and selected relations carry mechanical consequences. The architecture resembles proof checking in one limited respect: a powerful producer is separated from the smaller mechanism that decides whether its evidence is sufficient for admission.88. Talia Ringer et al., “QED at Large: A Survey of Engineering of Formally Verified Software,” Foundations and Trends in Programming Languages 5, nos. 2–3 (2019): 102–281, https://arxiv.org/abs/2003.06458. Three joins show the difference in increasing order of reach.

Property facts → verification tier → required checker → blocking consequence. Each invariant in the cross-service state-machine model carries declared property facts: how many concurrent lanes touch it, whether it is a safety or a liveness property, what state it ranges over. From those facts the model derives a verification tier — an invariant crossing multiple concurrent lanes demands exhaustive state-space exploration; a single-lane invariant demands a property test; a temporal obligation demands a model checker run. The tier is not advice: a check walks the model and fails when a mandated checker does not exist on disk. The supervisor does not decide, invariant by invariant, what evidence concurrency deserves. The representation decides, and the environment enforces the decision.

Service relationships × deployment wiring × runtime identity → an access-grant plan. The typed inter-service edge model declares which service calls which, over which plane, under which identity. Deployment machinery consumes those edges and derives the invoker grants and token audiences each service needs — and a coherence gate asserts at deploy time that the granted permissions match the modeled relationships. Least-privilege wiring is not a checklist a human re-audits; it is a projection of the service graph.

Mutation → stamp → derived account. Every mutating operation in the remediation core writes a provenance stamp — the verb, the pass, the visibility of the change — and the customer-facing changelog is derived from those stamps rather than authored beside them. A wiring validator identifies mutating paths that fail to stamp their work. The account of what the factory did to a document is therefore evidence, not narrative.

Modeling changed what the factory could know. Alignment changed what that knowledge could cause.

The Computation Graph: a Representation That Took Over Execution

The computation graph became the deepest of these joins over time. Its history shows why. The remediation core was organized early as passes — identifiable units of computation — and that decomposition proved durable. What stayed implicit was their composition: no pass declared, in one common vocabulary, what it produced, what it consumed, what it read only to decide whether to run, and how bounded its mutations were. Across a growing pipeline those locally reasonable freedoms interacted: work whose local complexity looked harmless accumulated quadratic behavior as passes re-derived state, and passes mutated the document through different mechanisms.

The response was to represent the composition well enough to reason about it. Deriving the missing producer–consumer relations from the passes raised a sharper question — consumer in what sense? Roughly two-thirds of the candidate relationships were not data flow at all but routing: a pass read a verdict only to decide whether it should run. So the passes acquired distinctions they had never expressed uniformly — produces, consumes, consumes-for-control — and the graph projects separate data-flow and control-gate relations from them. The same exercise classified every pass site by mutation kind: typed-patch producer, direct editor, or read-only. The counts did not say the direct editors were wrong; they showed where degrees of freedom remained. Figure 5.2-7 draws the progression.

The modeling history A top-to-bottom modeling history. DocAble began with a stable decomposition of remediation computations, so their node projection was meaningful, while their composition stayed implicit — carried by repeated traversal, shared-state reads, direct mutation, and special cases. As that composition accumulated global cost, modeling the relations exposed two previously implicit dimensions, shown as parallel columns: dependency semantics (DATA_FLOW, CONTROL_GATE, CROSS_SERVICE) and effect boundedness (TypedPatchProducer, DirectEditor, ReadOnly). Typed declarations then made the relationships projectable, a projected graph held by a blocking parity check. Analysis over that projected structure then motivated a redesign: an authoritative computation model that execution consumes to determine dependency order, with analytical views attaching by stable computation identity. Residual freedom remains. The progression runs from a partial model, to an analyzable one, to a model that participates directly in realization. The modeling history Stable computations pass decomposition — the nodes already project Composition stays implicit repeated traversal · shared-state reads · direct mutation · special cases Composition becomes consequential locally reasonable choices accumulate global cost model the relations two previously implicit dimensions Dependency semantics DATA_FLOW CONTROL_GATE CROSS_SERVICE Effect boundedness TypedPatchProducer DirectEditor ReadOnly Typed declarations each pass declares its IO facets and its effect kind Projected graph + blocking parity edges projected from the declarations; disagreement fails the build Analyze composition critical paths · boundedness · latency · coverage Authoritative computation model execution consumes modeled dependencies Analytical views attach by stable identity views join without one universal schema Residual freedom remains not every choice is determinized From a partial model, to an analyzable one, to a model that participates directly in realization.
Figure 5.2-7. The modeling history. DocAble acquired a stable decomposition of remediation computations before it acquired an explicit model of their composition. As locally reasonable choices accumulated consequential global cost, modeling exposed dependency semantics and effect boundedness. Typed declarations made those relationships analyzable and checkable; later redesign made explicit composition authoritative for execution, while separate analytical views continued to grow around stable computation identities. The progression is from a partial model to a richer and increasingly consequential one, not from no model to model.

Two further moves made the representation consequential. First, the relationships stopped being a second hand-written truth: the pass declarations carry the typed IO, the graph projects its edges from them, and a blocking parity check fails when declaration and projection disagree. Second, a later redesign made the model authoritative for execution itself — the execution machinery consumes the modeled dependencies to determine what becomes ready and in what order work may proceed, so composition no longer lives implicitly in the pass bodies. A representation introduced to understand the pipeline became part of the machinery that runs it. The analytical graph remains as a view over the modeled structure, carrying questions execution need not answer: critical paths, mutation boundedness, latency and cost, test coverage.

Explicit models changed not only what the system recorded but what questions the environment could ask of it. A lifecycle model made illegal transitions explicit. An ownership model reduced a concurrency protocol to the states and interleavings needed to ask whether two workers could hold the same claim; a small purpose-built checker could then exhaustively explore that finite state space. Eventual termination — a property over executions, not states — was written as small temporal specifications and checked mechanically; with recovery disabled, the checker produced exactly the stranded execution the property was meant to exclude. Checking the model created a second obligation: why believe a property established over the model holds of the implementation? Conformance tests exercised production operations against the executable model, and when one such test drifted red for days after production changed, the durable response was correspondence machinery — model-to-implementation checks selected automatically when modeled surfaces change. The factory then added machinery to keep the map true.

5.2.4 The Factory in Operation

With the representations and joins in place, the mature factory's production flow can be walked in one pass. What follows is everything the build accumulated, in motion. Engineering intent becomes a named work unit; a brief carries the task, its scope, and the context selected for it; realization fans out to agents working in isolated copies of the repository. Nothing an agent produces reaches the production branch directly. Each candidate change queues for a merge train that rebuilds, re-lints, and re-tests the actual candidate tree — the independent-evidence discipline a later episode in this subsection explains — and only work that survives that gate lands. Parallelism lives entirely on the realization side; admission stays serialized through regenerated evidence.

One operational week puts numbers on that flow (Figure 5.2-8). In the mature era's highest-throughput full week, the factory started 592 merge-train runs; 573 landed on main through the full integration gate and 10 aborted — a 96.8% completion rate — while 192 agent work spaces were retired after their work landed. Roughly six hundred serialized admission decisions in one week, each backed by regenerated evidence rather than the producer's claim, supervised by one person who read almost none of the underlying diffs. This is the concrete payoff of the accumulation the build section traced: the gate that admits hundreds of candidate changes without human diff-reading is assembled from the controls Figure 5.2-4 counted and the evidence machinery the episodes below explain. Two limits keep the figure honest: the week is the best-throughput full week (a representative mature week, not an average one), and the event log covers only the admission stage — per-task validation attempts, autonomous repair iterations, rollbacks, and human escalations were not retained as queryable events, so the funnel upstream of admission is not reconstructable from this record.

2026-09-25T11:43:38.382990 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
Figure 5.2-8. The Admission Funnel, One Mature Week. Merge-train runs in the mature era's highest-throughput full week: 592 started, 573 landed on main through the full integration gate, 10 aborted, 192 work spaces retired. Parallel realization fanned out; admission stayed serialized through regenerated evidence. The registry logs only this stage — upstream repair iterations are not retained as events.

Between realization and admission sits repair, and most failures die there, where they are cheap. An agent whose change fails a local check — the accumulated controls again, now doing their routine work — repairs it and tries again; a candidate that fails the merge train's gate is rejected without landing, and the work returns for another attempt. The engineer sees none of this. What reaches the engineer is the residue: failures that resist local repair, evidence the machinery cannot interpret, and — the consequential class — failures that implicate the engineering environment itself. Escalation, in this factory, is not a ticket queue; it is the moment a failure stops being an implementation problem and becomes an engineering one.

What Escalated Failures Became

Where did all this factory machinery come from? Not from a grand upfront design. The build section's control curve showed the count rising week over week; what follows reconstructs where selected pieces of it came from, one episode at a time. Governance conversion makes an empirical claim: some engineering episodes leave behind durable structure that later work inherits instead of paying for the same judgment again. What follows reconstructs selected DocAble incidents where that sequence — pressure, response, durable consequence — is actually visible, organized by what each failure converted into rather than by when it happened. It is not a claim that every structure arose from failure; the subsection ends with the ones that did not.

Figure 5.2-9 summarizes the twelve episodes by the durable structure each became. Eight ended in deterministic, blocking structure — an explicit state model, a typed seam with a ban, a lint floor, an evidence-regeneration gate. Two remained probabilistic or advisory: a decision rule and a routed judgment, kept where a deterministic check cannot carry the question. And two stopped at measurement — one because the evidence later earned an architectural change, one deliberately, because the evidence did not yet justify a gate. The distribution is the section's soft/hard axis made visible in one exhibit: alignment usually culminates in blocking enforcement here, and deliberately does not always. One caution: the episode set restates this chapter's own narrative in structured form — twelve episodes with approximate dates — not an independent census of every control the repository grew.

2026-09-25T11:43:38.796376 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
Figure 5.2-9. What Experience Became. The twelve governance-conversion episodes by the durable structure each failure became, coloured by enforcement class. Eight land deterministic and blocking; two stay probabilistic or advisory; two stop at measurement, one deliberately without a gate. Three ex-ante confirmations are excluded, keeping the conversion-versus-confirmation distinction; the exhibit restates the chapter's own narrative in structured form rather than censusing the repository independently.

An implicit lifecycle became an explicit behavioral model. The dominant production failure was a slow generation call, and the lifecycle around it was implicit — status written from scattered places, one terminal state masking many distinct failures, and requeue as the reflex recovery. Under a slow dependency that reflex became a storm: requeue, still slow, same timeout, requeue again — the recovery amplifying load exactly when the system was already hurting, every attempt billing its full cost with nothing marking it a retry. What became durable was the explicit lifecycle model from the previous subsection: named states and legal transitions, atomic claiming that separates permission to act from recovery leases, and no automatic requeue on ordinary failure. A later stale-owner sweep exposed the next missing distinction — slow is not dead — and progress evidence supplied it. Recovery stopped being a reflex and became a decision over explicit state and evidence.

Destructive library behavior became a typed seam and a ban. An early path used a convenient scripting-language library to slice a document. The operation worked visually and silently destroyed the internal tag structure carrying the accessibility the product exists to add — and nothing caught it. The response came in layers: a policy named the canonical library; a typed format model and API concentrated all mutation of that format through one seam; a build-time check refused raw access from elsewhere. Then the check itself failed — a stale repository root caused real findings to be reported as green — and the second-order incident forced the gate to be hardened too. The model stated the boundary; the check enforced it; the enforcement machinery itself still had to be maintained.

A cross-service assumption became a typed client and an explicit service graph. Another failure sent a local file path across a service boundary, where it became meaningless in the callee's process. The replacement client accepted file content but not a local path — the mistake became unrepresentable at the type level — and an explicit service graph represented which services could call which, the same graph the access-grant join now derives from. A later human review found a second problem: an optional direct-client path could bypass shared rate limiting, key rotation, failover, and cost tracking entirely. A protection boundary only governs the paths that cross it.

An underspecified instruction became a lint floor and named abstractions. By the time the system reached roughly 300,000 lines, regions built MVP-fast lacked coherent abstractions, so I told the agent to use typed Python. It did, in a sense: it used the type string for everything — stringly-typed, the word "type" satisfied to the letter and voided in spirit. "Use types" specified a surface property, not the domain abstractions the types were supposed to represent, and the agent took the cheapest satisfying path.†† My field notes on the period were less measured: "Claude is the fastest road to hell." It was the speed with which a weakly specified direction could turn into a large amount of coherent wrong structure. The durable response combined a floor with structure: every commodity lint the language ecosystems ship was turned on and driven to zero, each made blocking once the repository conformed; then, region by region, dense primitive code was routed through induced typed domain abstractions, with the checks promoted as each governed surface reached conformance. A month later the 300,000 lines were roughly 200,000, and future work inherited named abstractions, boundaries, and regression checks rather than a rewritten snapshot.

A recurring architectural ceiling became a decision rule. The service cost tens of dollars a day while idle on a cluster. The first response stayed inside the existing architecture — tune the autoscaler, reduce idle pods — and for weeks the agents made it better and better, with every improvement colliding with the same boundary: cold pods took too long to return. The repeated failure was information about the frame rather than the implementation inside it. The durable lesson was a decision rule — when repeated in-frame optimization collides with the same architectural ceiling, price the architectural change rather than commissioning another round of local optimization — and the eventual serverless migration encoded the decision in the system itself. A judgment becomes capital only when later work can inherit it; a lesson held in someone's head remains expensive to reuse.

An unmade decision became a decision made once. A merge automation kept stalling when two agents edited the same file, and the proposed one-flag fix — on conflict, take one side and drop the other — was signed off without asking what "drop the other side" meant. For four days it silently deleted real work from the main branch, the same commit heading landing again and again, each time a little smaller. The conflict handler contained a consequential design choice that had never been made explicitly: what information may be discarded when two changes overlap? At execution time the agent still had to choose, so it selected the locally expedient behavior. The resolver was changed to reconcile the conflicting files, and the destructive default was prohibited. The portable lesson is narrow: do not leave consequential semantics unresolved at a boundary where an autonomous actor must act.

"Done" became a claim requiring independent evidence. Two incidents, close together, made the rule non-negotiable. A lint gate hit a stray error, caught it in a catch-all that returned success, and let a run of commits land completely unlinted while the log said fine. A chain of pin tests passed green — because a botched merge had landed them onto a tree missing the very field they were supposed to check. Both reported done; neither was. A report of "done" describes the producer's claim, not the state of the repository, and the gap between the two is where a confident green hides. The durable response regenerated the evidence at the admission boundary: required tests and lints re-run against the actual candidate tree, gate failures made fail-loud — a control cannot translate its own uncertainty into success — and counts re-derived rather than accepted from the task report. The capital is not a standing habit of distrust; it is machinery that regenerates the evidence cheaply. That machinery is the merge train whose weekly funnel opened this subsection — the same gate, measured in operation.

An inadequate input representation became a routed judgment. An agent wrote a clean, typed, careful detector: if an image's pixels barely vary, skip the vision model and mark it decorative. Its own comment promised a false positive was "not possible" — real photos and charts vary twenty times more than the cutoff. The margin was true, and it was measured on the wrong world. Hand it a Rothko, or Malevich's black square, and the detector marks the painting decorative, and a blind student reading an art book gets silence where the subject should be. No threshold saves it, because a blank spacer and a low-variance artwork can be identical in the one quantity the code measures — "content-free" is a judgment about an image in context, not a fact about pixels. The cheap proxy lost the last word: ambiguous cases route to a richer semantic judgment, and the regression evidence covers the failure class — low-variance images that nevertheless carry content — rather than the first example that exposed it.

Forensic reconstruction became derived provenance. A corrupted document once took four cross-file hops to trace back to the pass that wrote it, and verification then found eleven production sites writing through a low-level path that skipped the provenance stamp — shipping empty changelog rows for a whole class of jobs while the code had changed plenty. The response is the provenance join described earlier. Its landing sequence matters: the wiring validator reported first, because the inherited tree contained many violations; a fix wave drained the gap; only then was the obligation enforced at the admission boundary. Representation made the obligation observable; migration made it feasible; gating enforced it.

A silent stall became measurement — and deliberately stopped there. A document's repair stalled in silence; the subprocess was killed by the job timeout after roughly eight minutes, the customer's progress bar frozen at thirty-eight percent, the budget already spent. The incident produced a provisional per-chunk worst-case model against the job budget — the seed kept as a timestamped datapoint in a file, not frozen as a constant in code — and instrumentation that can report a would-be breach. Figure 5.2-10 draws where the chain deliberately stops: no production gate rejects a job on this model today, because the bound is not calibrated strongly enough to justify enforcement, and a false refusal could cost more than continued observation. Alignment need not culminate in a blocking gate. The counterpart shows measurement earning action when its evidence is adequate: a deterministically measured per-request cold-start floor of 4,057 ms, joined to a deployment-topology model, drove a re-architecture that cut the floor to 109 ms. Same progression from representation to evidence; different verdicts, each proportionate to the evidence.

Modeled and Observed, Not Yet Binding — the measurement history that stops before the gate An engineering history, not a system model. A vertical chain. A silent hang and a bare timeout, which spoke only after the budget was spent, forced a measurement, then a per-chunk cost model, whose evidence lets the environment report a would-be breach. The chain ends at report — and then an enforcement node drawn as a ghost: a dashed box prefixed with a question mark and marked no gate yet, deliberately deferred. The model reports a predicted breach but does not refuse the work. The tail model, evidence, report rhymes with Chapter 3's ungated measurement. The lesson: a model can be left unbound on purpose — modeled and observed, reported, not yet enforced. A model left unbound on purpose: modeled, observed, reported, not enforced. timeout a bare kill, after the budget was spent measurement one incident, timestamped, provisional model per-chunk worst case vs the ceiling evidence a would-be breach, made observable report surfaced, not enforced ? enforcement [ no gate yet — deliberately deferred ] ? An engineering history, not a system model.
Figure 5.2-10. Modeled and Observed, Not Yet Binding. A timeout incident produced a measurement, a provisional per-chunk model, and reportable evidence. Production admission still does not depend on the model: its bound is not calibrated strongly enough to justify blocking. The terminal enforcement node is left open on purpose — representation and observation can mature before enforcement is earned. The open node is part of the evidence, not an omission.

And some conversions were almost mechanical. Invisible remediation artifacts were supposed to carry a reserved name prefix so a validator could always tell an inserted element from an authored one, and review kept catching insertions that skipped it. A prefix rule is entirely decidable, so the audit question became a blocking lint: the judgment moved from repeated inspection into the environment, where it fires the same way on every commit.

Not every structure was purchased by an incident. The backend became reactive and stateless early, which later made the serverless migration cheap; the document editor routed human gestures and automated remediation through one closed edit vocabulary, which later supported a second producer without a second mutation path; the unified detect-and-repair engine consolidated two scanning paths for a clean architectural reason, and a later audit found the boundary holding. Report those as confirmation, not conversion. Some assets were forged in incidents and hardened through recurrence; others were designed from known obligations and simply survived later pressure. The factory's machinery accumulated along both paths — but on neither path did it arrive as a plan drawn in advance.

5.2.5 The Supervisor's Job

What was the human doing while the factory accumulated? The answer changed, and the change is the section's payoff. The preceding subsections have each already recorded a piece of it. Work arrived as formalized intent rather than as code to read. System knowledge was consulted in representations rather than reconstructed from source. "Done" was established by evidence regenerated at the admission boundary rather than by the producer's claim. Failures reached the engineer as escalations rather than as diffs. The staircase below names the same transformation chronologically, from the supervisor's chair: each step moved some combination of system knowledge, realization work, evaluation, or admission out of direct human handling and into the engineered environment. Five rung labels name the changing center of gravity of my own work: co-coder, QA, review lead, tech lead, architect. What matters is what the engineered environment could carry at each stage. It is a chronology of DocAble, not a maturity model for other factories.

Co-coder: direct human inspection carried everything — I read the agent's changes, judged them, merged them. It worked while volume stayed inside one person's attention, and failed as an operating model once generation outran review. QA: evaluation was delegated before enforcement — a second agent audited another agent's work, episodically and probabilistically; its value was diagnostic, and its repeated decidable findings revealed questions that no longer needed a language model at all. Review lead: reviewers specialized by concern and region, and every finding that turned out to be mechanically decidable dropped into deterministic machinery, the probabilistic reviewers keeping the questions checks could not settle. Tech lead: representation made reasoning local — named zones externalized ownership and boundaries so an agent could reason inside one region without first recovering the whole system, and selected obligations expressed over the zone model were enforced. Architect: by the final rung I rarely read implementation directly. Architectural, behavioral, ownership, and other models supplied the system views for planning and review; agents used the same representations to localize changes. Declaration did not make the models true — model drift became its own failure class, which is why correspondence checking earned its place in the previous subsections. And at this rung the economics of restructuring inverted: a cross-format migration that touched 58 files and rewrote 5,000 lines took about 5 hours of agent time over a dinner break. The hard part was never touching fifty-eight files; it was knowing that fifty-eight files should be touched. Figure 5.2-11 draws the climb.

The delegation staircase — five ascending steps from Co-coder to Architect; delegated work climbs, the engineered environment accumulates, and residual human burden narrows but never vanishes A true staircase of five steps ascending from bottom-left to top-right. Step one at the bottom is Co-coder; step five at the summit is Architect. Each step card names a stage and shows, in blue, what is now delegated, and in green, what the engineered environment now carries. Along the left edge a green gauge thickens as the climb proceeds — representation and enforcement accumulate. Along the right edge a red trajectory narrows from wide at the base to a sliver at the summit: the residual human focus moves from reading every diff at the bottom to system-level strategy and residual semantic judgment at the top — the tactical burden shrinks but never reaches zero. Step one Co-coder: delegated is code generation; the environment carries nothing durable beyond an ordinary repository; the human reads every diff, holds the map, and admits the change. Step two QA: delegated is generation plus a second agent's audit; the environment carries a probabilistic, episodic, advisory reviewer; the human adjudicates findings. Step three Review lead: delegated is specialized review per region; the environment carries deterministic checks for the decidable findings; the human does the semantic review a check cannot settle. Step four Tech lead: delegated is bounded work inside one modeled zone; the environment carries named zones and boundary rules; the human holds the system map and strategy. Step five Architect: delegated is model-guided work across the whole system; the environment carries explicit models, traceability, and controls; the human is left with residual semantics and strategy. A failure in representation or enforcement can push work back down the staircase. HUMAN FOCUS strategy · residual semantics environment accumulates representation + enforcement residual human judgment narrows — never zero 5 · ARCHITECT explicit models carry system-level knowledge delegated — model-guided work across the system environment — explicit models · traceability · controls residual semantics + strategy 4 · TECH LEAD representation makes reasoning local delegated — bounded work inside one modeled zone environment — named zones · boundary rules system map + strategy 3 · REVIEW LEAD semantic residue after deterministic checks delegated — specialized review, per region environment — deterministic checks for the decidable findings semantic review 2 · QA evaluation delegated before enforcement delegated — generation + a second agent's audit environment — a probabilistic reviewer (episodic, advisory) adjudicate findings 1 · CO-CODER direct human inspection carries everything delegated — code generation environment — nothing durable beyond an ordinary repository read every diff · hold the map · admit the change more work carried by the engineered environment delegation climbs
Figure 5.2-11. The Delegation Staircase. As explicit representation and environmental enforcement accumulated, larger units of engineering work could be delegated. Representation moved system knowledge out of one person's head; enforcement moved repeatable admission decisions out of direct review. Human work moved from reading every diff toward system-level strategy and the residual semantic judgment the environment could not yet decide. A failure in either representation or enforcement could move work back down the staircase.

The staircase is a chronology, not a measurement, and the mining pass deliberately did not manufacture one. But one quantitative trace does sit behind the top rung (Figure 5.2-12). Across the window's 32,039 main commits, the commits that are not attributable to an agent or to integration automation — the explicit human-or-unknown residual — number 1,763, about 5.5%. That figure over-states human authorship: the residual bucket also holds early-history agent commits made before the attribution convention existed, which is why the first weeks of the series are nearly all residual and the convention's arrival in week 5 collapses it to single digits. And commit share is a different quantity from inspection share — no evidence records whether a human read a given change, so the trace bounds human-authored realization from above and says nothing directly about reading. Within those limits it is consistent with the qualitative claim the staircase makes: by the mature era, direct human realization was a rounding error against the fleet's.

2026-09-25T11:43:39.315163 image/svg+xml Matplotlib v3.10.9, https://matplotlib.org/
Figure 5.2-12. Weekly Commit Provenance as a Delegation Proxy. The explicit human-or-unknown residual is about 5.5% of the window's 32,039 main commits — an upper bound on human-authored realization, since early agent commits predating the attribution convention land in the residual. Commit share is a different quantity from how much the engineer read.

The staircase has two interacting supports, and delegation climbed only as far as they could carry it. Alignment moved enforcement of repeatable obligations out of direct human review: deterministic checks, permissions, boundaries, and gates enforced obligations I no longer inspected line by line. Modeling moved system knowledge out of my head: named zones and later explicit models let agents reason over larger engineering questions without reconstructing the whole repository each time. When representation was too weak, I fell back into reconstructing the map; when enforcement was too weak or unreliable, I fell back into reading diffs and verifying outcomes by hand. A miss sent the staircase backward.

Functionally, the top of the climb rewrites the supervisor's job description. As the MAGE layer strengthened, human attention moved away from supervising individual acts of realization and toward:

  • deciding which external pressures become engineering work;
  • choosing which engineering questions are consequential enough to represent;
  • constructing and revising the representations themselves;
  • deciding which obligations deserve enforcement, and at what strength;
  • interpreting anomalous evidence the machinery surfaces;
  • resolving genuinely ambiguous tradeoffs;
  • connecting representations to one another — the joins of the earlier subsection;
  • diagnosing failures in the engineered environment itself;
  • and deciding what experience should become durable structure.

Implementation gave way to engineering intent as the unit of supervised work; reading source gave way to explicit representations of consequential properties; the producer's claim gave way to independently regenerated evidence; routine inspection gave way to anomalies and escalations; per-diff review gave way to serialized admission decisions; isolated descriptions gave way to joins that carry consequences; and recurring local judgment was converted, episode by episode, into durable structure. MAGE afforded supervision by changing the objects through which supervision occurred. Not "MAGE automated supervision" — human judgment remains central, and every item on the list is judgment. But supervisory effort no longer scales linearly with implementation volume for any property that can be represented and mechanized, and that is the difference between the co-coder reading every diff and the architect reading almost none.

5.2.6 Economics

What did this factory cost, and what did the money buy? A factory's startup cost, its marginal change cost, and its carrying cost are different numbers — a separation §6.1 develops formally. DocAble puts observed numbers against that structure, and the accounting is deliberately simple (Table 5.2-3).

Table 5.2-3. Approximate observed cost of constructing and operating the DocAble factory through September 22, 2026. Claude token consumption is estimated; OpenAI token consumption is metered; costs are rounded. Roughly 90% of the observed investment was human engineering labor.
Factory inputQuantityCost
Human engineering6 months$100,000
Claude coding agents~5.0B tokens$4,800
OpenAI inference628M tokens$946
Google Cloud infrastructure—$4,886
Total~5.6B tokens$110,633

Over the six months I valued my engineering time at approximately $100,000. Four Claude Max subscriptions supplied the coding agents at a total subscription cost of $4,800; their aggregate consumption was heavy enough to be on the order of billions of tokens, and I estimate approximately 5 billion.‡‡ Anthropic does not expose Claude Max subscription usage as a fixed token allowance, so the Claude token quantity is an order-of-magnitude estimate rather than a metered measurement. The usage reported here was academic research use through individually purchased Claude Max subscriptions and Claude Code; Anthropic's Max plan explicitly includes Claude Code and is subject to Anthropic's usage limits. The subscriptions were used within their ordinary usage limits — at least insofar as my madness can be considered ordinary — rather than to bypass metered API charges. Metered OpenAI services consumed 628 million tokens for $946.45, and Google Cloud infrastructure cost $4,886.30. The composition is more informative than the total. Approximately 90% of the observed investment was human engineering labor, and the factory consumed on the order of 5.6 billion tokens of machine intelligence while direct agent and inference charges remained below $6,000. Commodity intelligence was not the expensive part of this factory. Engineering the factory was.

Even the infrastructure cost contains an instructive inefficiency. Roughly half of the Google Cloud expenditure resulted from an initially always-on Kubernetes deployment — the same idle-cluster cost whose repeated in-frame optimization became a decision rule earlier in this section — and the serverless migration that replaced it sharply reduced the carrying cost. The factory was itself an engineered system: its production economics changed as its architecture improved.

The $110,633 should not, however, be divided by the number of documents processed during these six months and reported as the cost of a document. Much of the expenditure constructed a durable production capability: the models, mechanisms, services, tests, and orchestration described in this chapter remain available to later work. Nor are the nearly two million lines of production and supporting software the return on the investment. Those artifacts are the factory purchased by the investment. The return lies outside the factory — documents remediated, remediation effort displaced, turnaround reduced, obligations satisfied — and assigning dollars to those outcomes would require a valuation model this case does not provide.

In those terms, the $110,633 is primarily an observation about Cstartup. The open empirical questions are now three different numbers: what the factory costs to keep, what the next substantial change costs, and what another document costs once the capability exists.

The Manual-Remediation Counterfactual

Was that a defensible investment? Directly observed construction and operating costs totaled approximately $111,000. To avoid false precision and allow generously for omitted costs, the analysis that follows treats the factory as a $200,000 investment.

The investment can be compared with the work the factory was built to perform. Purdue's public syllabus library listed 7,605 distinct courses for Fall 2026 across the university's three campuses — the count shown in the library when it was consulted on September 22, 2026, not an independently audited institutional statistic.99. Purdue University, “Syllabus Library,” 2026, https://purdue.simplesyllabus.com/en-US/syllabus-library. A conservative order-of-magnitude estimate of 10,000 distinct courses across an academic year follows.

Suppose manually reviewing and remediating one course's instructional materials takes one person-week: 40 hours to inspect and repair, say, thirty slide decks and fifty readings, at an institutional labor cost of $100 per hour including benefits and overhead.

40 hours×$100/hour=$4,000 per course

These are assumptions, not findings, and they are deliberately conservative: this section opened by estimating 304 hours for one course at slower per-document rates, and a colleague reported spending a week on a single PDF. Choosing the low figure biases every comparison that follows against the factory.

Under these assumptions, comprehensive manual remediation of 10,000 courses represents approximately $40 million of labor-equivalent work:

10,000 courses×$4,000/course=$40,000,000

That is a one-time figure, not an annual one. The existing instructional corpus must be remediated once; the recurring workload is only new and changed material, so the manual counterfactual has the same shape as the factory itself:

Cmanual=Cinitial corpus+Congoing changes

The ongoing term depends on the rate at which instructional materials are added or changed, and no estimate of it is offered here.

Purdue would not necessarily spend $40 million remediating every course manually — no institution was going to hire four hundred person-years of remediation labor. The calculation establishes a counterfactual unit cost against which productive capacity can be evaluated. Extending the same arithmetic to only 100 Purdue-scale institutions yields approximately $4 billion of labor-equivalent work; the close of this section returns to that scenario.

What the Rulemaking Record Says the Alternative Is

These figures measure the scale of the work under a manual-production counterfactual; they are not a market forecast. Public educational institutions do not have unlimited budgets for accessibility, and work does not become funded merely because a legal or engineering obligation exists. At sufficiently high unit cost, the likely alternative to comprehensive remediation is not spending whatever comprehensive remediation costs. It is doing less.

The rulemaking record documents that response, in two complementary observations.3 In the Department's account, a Congressman "emphasized that current technology, including generative AI (artificial intelligence), cannot reliably automate the remediation of STEM materials at scale, and human oversight is required to ensure accessibility." Elementary and secondary education advocacy associations separately emphasized that many school districts have "limited financial and staff resources available for compliance," and one association warned that the compliance schedule risked overwhelming schools, which could cause them to attempt "rapid, procedural box-checking" rather than "thoughtful and sustainable implementation efforts that would maximize the goals and benefits of the rule."

Together, these observations describe an engineering problem with no satisfactory production system. Manual remediation is expensive at the required scale; existing automation was judged insufficiently reliable for difficult material; and resource-constrained institutions risk substituting procedural compliance for the accessibility outcome the rule is intended to produce. Accessibility is an outcome, not a checklist. A process can record that prescribed checks occurred while users continue to encounter inaccessible material — procedural compliance is not the same thing as achieving the outcome the obligation exists to produce.

DocAble aims into the gap between those two statements. The economic opportunity is not $4 billion of existing expenditure waiting to be displaced; institutions could not purchase comprehensive manual remediation at that scale, so that spending does not exist. The opportunity is to change the production function: to combine inexpensive machine realization with sufficient models, validation, and human oversight that reliable remediation becomes economical at the required scale — an outcome that neither comprehensive manual work nor the automation the record judged insufficient could economically provide.§§ This is also why I regard the present transition with considerably more excitement than apprehension. Commodity intelligence does not merely make existing engineering work cheaper. It can make previously impractical engineering economically tractable. Problems that could not justify enough skilled human effort can become problems that a skilled engineer can now afford to attack. The risks created by that capability deserve serious engineering attention — the subject of much of this book — but so does the extraordinary expansion of what engineers can attempt.

The pattern extends beyond accessibility. The value of a new production technology cannot always be estimated from historical spending on the work it replaces, because historical spending is itself constrained by the old production function. When an outcome is too expensive, organizations ration it, accept degraded substitutes, defer it, or leave it unproduced. A sufficiently large change in production economics can therefore create a practical market rather than merely capture an existing one.

This is also why §6.1 leaves the value side of the factory model abstract. Had the model valued "documents processed" or "courses remediated," it would miss precisely this phenomenon. The value is the accessibility users actually experience; checklists and conformance measurements are models of — and evidence for — that outcome, chosen because they make it available for decisions.

Break-Even, Marginal Cost, and Replication

What must the changed production function deliver before the investment is recovered? Under the assumptions above, even the deliberately inflated $200,000 construction cost is equivalent to only fifty manually remediated courses:

$200,000$4,000/course=50 course-equivalents

Fifty courses are approximately 0.5% of the estimated annual course population: the factory need successfully substitute for one course in every two hundred before the value of displaced manual remediation equals the generously estimated cost of constructing the factory. A break-even threshold has an advantage a raw ratio lacks: it can be falsified. And the conclusion does not depend on the scale estimates. Even substantial error in the opportunity — by an order of magnitude — leaves the factory economics interesting.

The progression this section has followed — factory investment, manual counterfactual, break-even volume — has a fourth step: marginal production cost. The early evidence is a single pilot, reported as observed rather than audited: remediating roughly ten slide decks totaling about five hundred slides cost approximately $10 in model spend plus about ninety minutes of human review. The $200,000 was not spent to remediate the first fifty courses cheaply. It was spent to construct productive capacity whose next course can be remediated far below the $4,000 manual counterfactual — the marginal question §6.1 treats as the important one.

One scenario remains. Suppose 100 U.S. public universities operate instructional estates roughly comparable in scale to Purdue's — about 6% of the 1,570 public degree-granting institutions counted in 2022–23,5 stated as an explicit scenario rather than a claim about the actual distribution. Table 5.2-4 puts the four scales side by side.

Table 5.2-4. The remediation task at four scales under the stated manual-work counterfactual: one course, the factory's break-even volume, a Purdue-scale institution, and an explicit 100-institution scenario. The rows share one assumed unit cost — $4,000 of manual labor per course.
ScaleCoursesManual remediation equivalent
One course1$4,000
Factory break-even50$200,000
Purdue-scale institution~10,000~$40M
100 Purdue-scale public universities~1M~$4B

These figures are not estimates of expenditures universities will actually make. They characterize the scale of the remediation task under a stated manual-work counterfactual. The initial corpus must be remediated once; subsequent costs depend on the rate at which instructional materials are added or changed.

The last row is where software-factory economics diverge from the physical analogy. DocAble's construction cost does not multiply by 100 universities: Purdue does not build a remediation plant and then Stanford build another. Subject to deployment, integration, and operating costs, the software capability replicates, and three amortizations stack:

factory construction→courses remediated→institutions served

The first fifty course-equivalents can repay the generously rounded construction cost; Purdue's corpus alone is two orders of magnitude past that threshold; and the same software can serve further institutions without repeating the six months of factory construction. This is the distinction §6.1 draws between startup cost, marginal production cost, and the near-zero replication cost of software capital, demonstrated on one factory's books.

Keep three numbers distinct. Roughly $4 billion of remediation-equivalent work is the engineering opportunity. What institutions would actually pay is the addressable market. What DocAble captures is a business question. MAGE needs only the first.

5.2.7 Liability

Keeping the case honest requires the other ledger column, because everything this section has praised also has a failure mode. One model named a code symbol that was later moved and renamed. Nothing invalidated the reference; the model continued to look authoritative, agents continued to reason from it, and the failure surfaced days later during deployment. A later audit showed the stale pointer was a class, not a one-off — re-running closed work against current code found genuine model-to-code drifts a green definition of done had missed. That incident is why correspondence checking exists, and it generalizes.

A hundred twenty-nine models can drift, and the correspondence machinery that catches that drift must itself be maintained. Controls can conflict, and resolving a collision between two well-intentioned gates is supervisory work no model retires. Evidence goes stale: a baseline recorded under last quarter's workload quietly stops meaning what its readers assume. Gates encode the assumptions of the incident that created them, and an obsolete assumption enforced mechanically is worse than one merely remembered. The support apparatus — three times the production source — must itself be operated, debugged, and paid for in attention. The structures that carry decisions forward must themselves be maintained, and some will stop being worth their upkeep: not all of what this factory accumulated will survive its next architectural turn. The support ratio bought supervision; it did not buy it for free.

The evidence has clear limits. This is one longitudinal production case led by one primary engineer directing an agent fleet, using one contemporary model ecosystem, without a controlled comparison against another engineering process. It cannot establish that another organization will encounter the same failures, build the same mechanisms, or achieve the same economics — and checker scores do not establish complete end-user accessibility. What the case can establish is mechanism and sequence: these pressures occurred, these responses followed, and selected structures persisted long enough to be exercised under continuing change.

DocAble shows one deliberately engineered response to abundant realization. The factory therefore has to be evaluated as a production system, including the structures that direct and supervise realization, rather than by agent capability alone. The next question is whether independently developed agentic software factories exhibit the same pressures — and that question does not depend on anyone else having heard of MAGE.

Works Cited

  1. National Federation of the Blind. “Creating Nonvisually Accessible Documents.” 2014. https://nfb.org/images/nfb/products_technology/creating_accessible_documents.docx.
  2. U.S. Department of Justice, “Nondiscrimination on the Basis of Disability; Accessibility of Web Information and Services of State and Local Government Entities,” 2024.
  3. U.S. Department of Justice. “Extension of Compliance Dates for Nondiscrimination on the Basis of Disability; Accessibility of Web Information and Services of State and Local Government Entities.” Federal Register 91 (2026): 20902–12. https://www.federalregister.gov/documents/2026/04/20/2026-07663/extension-of-compliance-dates-for-nondiscrimination-on-the-basis-of-disability-accessibility-of-web.
  4. National Center for Education Statistics. “Table 3. Number of Operating Public Elementary and Secondary Schools, By School Type, Charter, And State or Jurisdiction: School Year 2023–24.” 2025. https://nces.ed.gov/ccd/tables/202324_summary_3.asp.
  5. National Center for Education Statistics. “Table 317.20. Degree-Granting Postsecondary Institutions, By Control and Classification of Institution and State or Jurisdiction: Academic Year 2022–23.” 2024. https://nces.ed.gov/programs/digest/d23/tables/dt23_317.20.asp.
  6. UsableNet. “ADA Accessibility Lawsuit Tracker.” 2026. https://info.usablenet.com/ada-website-compliance-lawsuit-tracker.
  7. Carlson, Laura. “Higher Ed Accessibility Lawsuits, Complaints, And Settlements.” 2026. https://www.d.umn.edu/~lcarlson/atteam/lawsuits.html.
  8. Ringer, Talia, Karl Palmskog, Ilya Sergey, Milos Gligoric, and Zachary Tatlock. “QED at Large: A Survey of Engineering of Formally Verified Software.” Foundations and Trends in Programming Languages 5, nos. 2–3 (2019): 102–281. https://arxiv.org/abs/2003.06458.
  9. Purdue University. “Syllabus Library.” 2026. https://purdue.simplesyllabus.com/en-US/syllabus-library.
© James C. Davis, 2026–present