Industry case studies
MAGE did not run at this company. This page reads an independent practitioner report through MAGE’s vocabulary, to see how cleanly an outside system maps onto the theory. Every correspondence below is the book’s reading of the source, never a claim the company makes about MAGE.
These cases are practitioner reports, not replications of MAGE and not causal tests. A correspondence cell states how cleanly the source's described behavior maps onto a MAGE construct as the BOOK reads it — never a claim Cloudflare (or any site) makes about MAGE, and never proof. not-described means the SOURCE is silent; absence from a report is NOT evidence the company lacks the practice.
Meet the case
Uber is the case where the economics came first. Faced with agents already writing most of its code, Uber treated the surrounding environment—not the model—as the thing to engineer: real work becomes benchmarks, models are chosen on cost-per-completed-task against those benchmarks, bounded subwork is routed to cheaper models, context and tools are compiled down so the reasoner spends fewer tokens, recurring workflows harden into reusable skills, and managed agents are judged in outcome units like cost per merged PR. The book reads the case as the clearest factory-scale view of the agentic environment as an object of continuous measurement and optimization—recognizably MAGE-like in everything except the modeling tier, where the richest structure stays a knowledge/context graph rather than an executable model of the system's behavior.
Distinctive starting point: software-factory economics at scale
What the case shows
Uber engineers the economics and reasoning-horizon of the environment around replaceable models, reaching a knowledge/context tier without an executable model of the system's behavior.
What the agents do
Managed agents span code review (uReview, on all PRs), self-healing CI failures, completing end-to-end PRs with visual validation, triaging on-call alerts, debugging incoming bugs, and a range of code-maintenance tasks, with human reviews and escalations. Local and cloud agents author most PRs; a primary model decomposes and evaluates while cheaper subagents execute bounded work.
Setting: org-wide · brownfield
Scale, as the source reports it:
- more than 70% of pull requests attributed to local or cloud agents; code output per engineer roughly doubled year over year
- 3,600+ agent skills across the SDLC; 30K+ skill executions/day; 25 pre-built code-mode skills
- weekly active agentic users 7x and weekly agentic requests 9.4x (Feb-Aug 2026); cost per 1,000 model requests down ~34% and cost per session down 52% from their peaks
- AI Context Graph: 24M nodes / 80M edges (86 node types, 117 edge types) integrating 30+ internal systems; a graph-grounded task ran in 38s vs 20+ minutes ungrounded
- code-mode collapses N model turns into one subprocess script (>90% token savings on bulk workflows, 55-71% on common SQL); 1,000+ MCP tools projected as CLI commands to avoid 50-70K tokens of schema overhead
The engineered environment
Object territory
source-code incident-reports design-docs
Representations
architecture-model tests informal-knowledge
Mechanisms observed
llm-reviewer retrieval-layer constrained-api merge-gate provenance
Where authority sits
Managed agents act autonomously on review, CI self-healing, alert triage, and maintenance with human reviews/escalations; PR merges and tier upgrades stay human, and model selection allows manual overrides. The primary model holds decomposition/evaluation authority; subagents are defaulted to a weaker, cheaper model. Authority is engineered per workload, not set in the prompt.
Mapping into MAGE
Each row is one MAGE construct, the strength of the correspondence, and how the book reads the source against it. The note is the book’s reading; the strength is not a score.
| MAGE construct | Correspondence | How the book reads the case |
|---|---|---|
| Alignment | ◐ partial | the book reads outcome-denominated gates, human merge/tier authority, and per-workload guardrails as Alignment behavior, though authority-without-the-author is less emphasized than at the policy-first or fleet-first sites |
| Modeling | ◐ partial | the book reads the AI Context Graph and real-work benchmarks as rich KNOWLEDGE representations while the governed SYSTEM'S behavior is not modeled as a primary reasoning surface (modeling ceiling) |
| Knowledge rep | ✓ strong | the book reads the 24M-node AI Context Graph plus 3,600+ reusable skills as a strong externalization of engineering knowledge into queryable, inheritable structure |
| Bootstrap | ◐ partial | the book reads skills, standard agent configs, and benchmarks as reusable E(0) that later work drops into, short of a full prior-standards estate |
| Conversion | ◐ partial | the book reads recurring-workflow->skill packaging and benchmark-driven optimization as capital formation; the source shows accumulation more than a measured failure->control->recurrence-retired loop |
| Determinization | ◐ partial | the book reads code-mode's deterministic subprocess orchestration and deterministic model/tool routing as moving decidable work off per-call inference |
| Reasoning horizon | ✓ strong | the book reads code-mode, CLI tool projection, graph grounding, compaction, and caching as the sample's most explicit case of the environment performing reasoning that would otherwise burn model context |
| Engineer's seat | ◐ partial | the book reads human ownership of merges, tier upgrades, objectives, and overrides as authority held outside the agent, alongside heavy delegated autonomy |
| Graduated governance | ◐ partial | the book reads primary-vs-subagent routing, per-workload model selection, and human-gated tier upgrades as graduated per-work authority, short of an explicit advisory->enforced promotion lifecycle |
The theory the case appears to hold
The book reads Uber as treating the engineered environment itself as the object of continuous optimization: measure the factory in outcome units, benchmark models on real work, route each workload to the most cost-efficient model, compile context and tools down so the reasoner spends fewer tokens, and harden recurring workflows into skills the whole fleet inherits. Reliability comes from the environment's economics and measurement, not from any single model.
What the case adds to MAGE
MAGE reads Uber's cost equation, benchmark-driven routing, code-mode, and skills as engineering the reasoning-horizon and the economics of the environment around a replaceable reasoner: the environment, not the model, is tuned to convert abundant intelligence into durable throughput at bounded cost.
environment-first-agent-economics environment-improvement-can-target-context-selection the-context-router-can-be-deterministic skills-encode-roles-not-procedures
What MAGE adds that this case does not reach
MAGE is a theory of the engineering ENVIRONMENT ITSELF as the object of engineering: everyone else engineers an agent, a runtime, a policy engine, or a model; MAGE engineers the governed environment in which commodity intelligence operates.
None of the external cases we examined describes the following machinery in the generalized form MAGE does. Not 'nobody in industry has ever done this.'
Honest bounds
The limitations the analysis records, and the falsifiable hypotheses the case bears on:
S e l f - r e p o r t e d o p e r a t i o n a l m e t r i c s , n o t a c a u s a l t e s t ; t h e a c c o u n t i s s t r o n g o n c o n t e x t , s k i l l s , r o u t i n g , b e n c h m a r k s , a n d c o s t / q u a l i t y m e a s u r e m e n t b u t d e s c r i b e s n o e x e c u t a b l e b e h a v i o r a l , s c e n a r i o , o r i n v a r i a n t m o d e l , a n d n o m o d e l < - > i m p l e m e n t a t i o n c o r r e s p o n d e n c e c h e c k , a s a p r i m a r y s o f t w a r e - r e a s o n i n g s u r f a c e . I t s s o p h i s t i c a t i o n t h e r e f o r e s t r e n g t h e n s r a t h e r t h a n d i s s o l v e s t h e m o d e l i n g c e i l i n g .
H4-representation-leverage H4a-representation-efficiency H6-oversight-amortization H3-mechanized-assurance
Source
| Field | Value |
|---|---|
| Citation | medisetty2026factory |
| Source type | practitioner-report |
| Independence | independent |
| Account type | retrospective |
| Evidence horizon | months |
| Author role | engineering (Distinguished Engineer, engineering productivity; internal practice account) |
Theory coverage at a glance
Where the source’s described behavior maps onto each MAGE construct: ✓ strong · ◐ partial · ~ tension · ✗ counterexample · — not described.
| MAGE construct | Correspondence |
|---|---|
| Alignment | ◐ partial |
| Modeling | ◐ partial |
| Knowledge rep | ✓ strong |
| Bootstrap | ◐ partial |
| Conversion | ◐ partial |
| Determinization | ◐ partial |
| Reasoning horizon | ✓ strong |
| Engineer's seat | ◐ partial |
| Graduated governance | ◐ partial |