Duality Lab — James C. Davis

Industry case studies

ReconstructionUber · software / ride-hailing + logistics platform

MAGE did not run at this company. This page reads an independent practitioner report through MAGE’s vocabulary, to see how cleanly an outside system maps onto the theory. Every correspondence below is the book’s reading of the source, never a claim the company makes about MAGE.

These cases are practitioner reports, not replications of MAGE and not causal tests. A correspondence cell states how cleanly the source's described behavior maps onto a MAGE construct as the BOOK reads it — never a claim Cloudflare (or any site) makes about MAGE, and never proof. not-described means the SOURCE is silent; absence from a report is NOT evidence the company lacks the practice.

Meet the case

Uber is the case where the economics came first. Faced with agents already writing most of its code, Uber treated the surrounding environment—not the model—as the thing to engineer: real work becomes benchmarks, models are chosen on cost-per-completed-task against those benchmarks, bounded subwork is routed to cheaper models, context and tools are compiled down so the reasoner spends fewer tokens, recurring workflows harden into reusable skills, and managed agents are judged in outcome units like cost per merged PR. The book reads the case as the clearest factory-scale view of the agentic environment as an object of continuous measurement and optimization—recognizably MAGE-like in everything except the modeling tier, where the richest structure stays a knowledge/context graph rather than an executable model of the system's behavior.

Distinctive starting point: software-factory economics at scale

What the case shows

Uber engineers the economics and reasoning-horizon of the environment around replaceable models, reaching a knowledge/context tier without an executable model of the system's behavior.

What the agents do

Managed agents span code review (uReview, on all PRs), self-healing CI failures, completing end-to-end PRs with visual validation, triaging on-call alerts, debugging incoming bugs, and a range of code-maintenance tasks, with human reviews and escalations. Local and cloud agents author most PRs; a primary model decomposes and evaluates while cheaper subagents execute bounded work.

Setting: org-wide · brownfield

Scale, as the source reports it:

The engineered environment

Object territory

source-code incident-reports design-docs

Representations

architecture-model tests informal-knowledge

Mechanisms observed

llm-reviewer retrieval-layer constrained-api merge-gate provenance

Where authority sits

Managed agents act autonomously on review, CI self-healing, alert triage, and maintenance with human reviews/escalations; PR merges and tier upgrades stay human, and model selection allows manual overrides. The primary model holds decomposition/evaluation authority; subagents are defaulted to a weaker, cheaper model. Authority is engineered per workload, not set in the prompt.

Mapping into MAGE

Each row is one MAGE construct, the strength of the correspondence, and how the book reads the source against it. The note is the book’s reading; the strength is not a score.

MAGE constructCorrespondenceHow the book reads the case
Alignment◐ partialthe book reads outcome-denominated gates, human merge/tier authority, and per-workload guardrails as Alignment behavior, though authority-without-the-author is less emphasized than at the policy-first or fleet-first sites
Modeling◐ partialthe book reads the AI Context Graph and real-work benchmarks as rich KNOWLEDGE representations while the governed SYSTEM'S behavior is not modeled as a primary reasoning surface (modeling ceiling)
Knowledge rep✓ strongthe book reads the 24M-node AI Context Graph plus 3,600+ reusable skills as a strong externalization of engineering knowledge into queryable, inheritable structure
Bootstrap◐ partialthe book reads skills, standard agent configs, and benchmarks as reusable E(0) that later work drops into, short of a full prior-standards estate
Conversion◐ partialthe book reads recurring-workflow->skill packaging and benchmark-driven optimization as capital formation; the source shows accumulation more than a measured failure->control->recurrence-retired loop
Determinization◐ partialthe book reads code-mode's deterministic subprocess orchestration and deterministic model/tool routing as moving decidable work off per-call inference
Reasoning horizon✓ strongthe book reads code-mode, CLI tool projection, graph grounding, compaction, and caching as the sample's most explicit case of the environment performing reasoning that would otherwise burn model context
Engineer's seat◐ partialthe book reads human ownership of merges, tier upgrades, objectives, and overrides as authority held outside the agent, alongside heavy delegated autonomy
Graduated governance◐ partialthe book reads primary-vs-subagent routing, per-workload model selection, and human-gated tier upgrades as graduated per-work authority, short of an explicit advisory->enforced promotion lifecycle

The theory the case appears to hold

The book reads Uber as treating the engineered environment itself as the object of continuous optimization: measure the factory in outcome units, benchmark models on real work, route each workload to the most cost-efficient model, compile context and tools down so the reasoner spends fewer tokens, and harden recurring workflows into skills the whole fleet inherits. Reliability comes from the environment's economics and measurement, not from any single model.

What the case adds to MAGE

MAGE reads Uber's cost equation, benchmark-driven routing, code-mode, and skills as engineering the reasoning-horizon and the economics of the environment around a replaceable reasoner: the environment, not the model, is tuned to convert abundant intelligence into durable throughput at bounded cost.

environment-first-agent-economics environment-improvement-can-target-context-selection the-context-router-can-be-deterministic skills-encode-roles-not-procedures

What MAGE adds that this case does not reach

MAGE is a theory of the engineering ENVIRONMENT ITSELF as the object of engineering: everyone else engineers an agent, a runtime, a policy engine, or a model; MAGE engineers the governed environment in which commodity intelligence operates.

None of the external cases we examined describes the following machinery in the generalized form MAGE does. Not 'nobody in industry has ever done this.'

Honest bounds

The limitations the analysis records, and the falsifiable hypotheses the case bears on:

S e l f - r e p o r t e d o p e r a t i o n a l m e t r i c s , n o t a c a u s a l t e s t ; t h e a c c o u n t i s s t r o n g o n c o n t e x t , s k i l l s , r o u t i n g , b e n c h m a r k s , a n d c o s t / q u a l i t y m e a s u r e m e n t b u t d e s c r i b e s n o e x e c u t a b l e b e h a v i o r a l , s c e n a r i o , o r i n v a r i a n t m o d e l , a n d n o m o d e l < - > i m p l e m e n t a t i o n c o r r e s p o n d e n c e c h e c k , a s a p r i m a r y s o f t w a r e - r e a s o n i n g s u r f a c e . I t s s o p h i s t i c a t i o n t h e r e f o r e s t r e n g t h e n s r a t h e r t h a n d i s s o l v e s t h e m o d e l i n g c e i l i n g .

H4-representation-leverage H4a-representation-efficiency H6-oversight-amortization H3-mechanized-assurance

Source

FieldValue
Citationmedisetty2026factory
Source typepractitioner-report
Independenceindependent
Account typeretrospective
Evidence horizonmonths
Author roleengineering (Distinguished Engineer, engineering productivity; internal practice account)

Theory coverage at a glance

Where the source’s described behavior maps onto each MAGE construct: ✓ strong · ◐ partial · ~ tension · ✗ counterexample · — not described.

MAGE constructCorrespondence
Alignment◐ partial
Modeling◐ partial
Knowledge rep✓ strong
Bootstrap◐ partial
Conversion◐ partial
Determinization◐ partial
Reasoning horizon✓ strong
Engineer's seat◐ partial
Graduated governance◐ partial