Industry case studies
MAGE did not run at this company. This page reads an independent practitioner report through MAGE’s vocabulary, to see how cleanly an outside system maps onto the theory. Every correspondence below is the book’s reading of the source, never a claim the company makes about MAGE.
These cases are practitioner reports, not replications of MAGE and not causal tests. A correspondence cell states how cleanly the source's described behavior maps onto a MAGE construct as the BOOK reads it — never a claim Cloudflare (or any site) makes about MAGE, and never proof. not-described means the SOURCE is silent; absence from a report is NOT evidence the company lacks the practice.
Meet the case
Meet Shopify through River, an agent you summon from a public Slack channel to read code, run tests, query the data warehouse, and open a pull request. River is public on purpose: a lesson learned in a private session helps no one else, so every session becomes a searchable artifact the company mines into skills and defaults. Underneath sits the deliberate part, a single work surface that centralizes context and reproducible execution, where each new agent is a profile over the shared substrate rather than a platform of its own. Shopify's distinctive habit is organizational rather than technical, keeping the work visible so one engineer's experience becomes everyone's, on a short evidence horizon with outcome numbers the company reports itself.
Distinctive starting point: shared agent-ready environment + organizational learning
What the case shows
Shopify made the engineering environment itself - a shared monorepo, reproducible Nix execution, a public agent corpus, and a common agent runtime - into compounding organizational memory, giving MAGE its strongest external instance of engineering capital and the refinement that experience must become observable before it can compound.
What the agents do
River is summoned from public Slack to do open-ended knowledge + implementation work over a durable session (read code, run tests, query the data warehouse, inspect production traces, open PRs, occasionally challenge a plan); multiplayer (a thread can attract another engineer and a new constraint without losing history); one profile among many planned. Humans merge.
Setting: org-wide · brownfield
Scale, as the source reports it:
- 59,918 River sessions across 5,170 Slack channels touching 7,000+ employees; 3,536 merged River-coauthored PRs in 30 days
- ~1 in 8 company-wide merged PRs coauthored by River; median session 19 min / 50 tool calls
- Quick: >50,000 internal sites created by >50% of the company
- Autoresearch applied across 40+ metrics (Shopify-reported deployment gains, not independently measured); Roast analyzed thousands of test files
The engineered environment
Object territory
source-code design-docs
Representations
informal-knowledge spec tests
Mechanisms observed
merge-gate deterministic-lint constrained-api retrieval-layer provenance stable-identity
Where authority sits
Reasoning autonomy inside environmentally-bounded execution: sandbox policy carried in profiles, the harness runs outside the sandbox, disposable sandboxes, Roast restricts file-writing/command execution through tool interfaces. Boundaries live outside the agent; merge stays human. [credentials-via-proxy UNVERIFIED in the fetch - re-verify for the prose wave]
Mapping into MAGE
Each row is one MAGE construct, the strength of the correspondence, and how the book reads the source against it. The note is the book’s reading; the strength is not a score.
| MAGE construct | Correspondence | How the book reads the case |
|---|---|---|
| Alignment | ✓ strong | the book reads sandbox policy in profiles, the harness-outside-sandbox split, disposable sandboxes, Nix reproducibility, and Roast's deterministic steps + tool restrictions as reasoning autonomy inside bounded execution |
| Modeling | ◐ partial | the book reads a broad representation surface (skills, conventions, intent, Nix) spanning the ladder, but with no invariant-bearing system/process models - Nix and Roast are the executable exceptions |
| Knowledge rep | ◐ partial | the book reads World + public-session mining into skills/conventions/defaults as real but SOFTER externalization — organizational knowledge retrieved on demand, mostly informal-knowledge rather than a structured system representation (a named piece missing -> partial) |
| Bootstrap | ✓ strong | the book reads the World monorepo + Nix + CI + merge queues stood up in 2024 (predating River, explicitly as an agent substrate) as an E(0) built ex-ante |
| Conversion | ✓ strong | the book reads private-learning->public-only, free-roaming->Roast, proliferation->Aquifer, and misuse->rate-limit as repeated failure->structure loops (the converted artifact often soft) - a strong-vs-partial contrast with Cloudflare, whose incident->RFC->rule loop is described but shows no measured recurrence-drop (partial) |
| Determinization | ✓ strong | the book reads Roast/Boba as putting deterministic machinery (sed cleanup, Sorbet autocorrect, tests) on the decidable residue and probabilistic judgment only on the rest |
| Reasoning horizon | ✓ strong | the book reads on-demand skills, the durable-session != persistent-context-window split, World's context centralization, and Roast replay as reasoning-horizon management |
| Engineer's seat | ✓ strong | the book reads humans supplying constraints/redirects, curating the environment, mining the corpus into skills, and owning merge + the platform boundary ('we've gotten good at saying no') as the engineer's seat |
| Graduated governance | ◐ partial | the book reads no explicit approved->enforced promotion lifecycle (Cloudflare's is stronger); the soft-knowledge accretion (conversation->skill->default) and Roast's structure-on-demand are graduated only in spirit |
The theory the case appears to hold
The book reads Shopify as holding that autonomous agents become organizationally useful when the environment is made common, reproducible, observable, and shared: consolidate the work surface, encode institutional knowledge in files/skills loaded at task time, make execution reproducible, separate durable sessions from disposable reasoning/execution, bound work in sandbox policies, make each new agent a profile over a shared platform, keep work public so local experience becomes shared knowledge, and combine probabilistic reasoning with deterministic steps for repeatable processes.
What the case adds to MAGE
MAGE connects Shopify's separately-named World / Aquifer / River / Profiles / Skills / Roast / Quick as representation + soft-conditioning + constraint + validation + feedback inside ONE governed environment, and names the structural-diagnosis question Shopify performs but never states.
publicness-precondition-for-compounding durable-session-as-model-independent-asset environment-first-agent-economics soft-governance-compounds multiplayer-as-environment-property
What MAGE adds that this case does not reach
MAGE is a theory of the engineering ENVIRONMENT ITSELF as the object of engineering: everyone else engineers an agent, a runtime, a policy engine, or a model; MAGE engineers the governed environment in which commodity intelligence operates.
None of the external cases we examined describes the following machinery in the generalized form MAGE does. Not 'nobody in industry has ever done this.'
Honest bounds
The limitations the analysis records, and the falsifiable hypotheses the case bears on:
no-invariant-bearing-system-model-described no-generalized-model-territory-drift-gate no-before-after-causal-performance-data short-evidence-horizon no-false-positive-or-drift-rates-reported four-load-bearing-details-unverified-in-fetch
H2-governance-moderation H3-mechanized-assurance H6-oversight-amortization H8-learning-propagation
Source
| Field | Value |
|---|---|
| Citation | moreno2026river |
| Source type | practitioner-report |
| Independence | independent |
| Account type | retrospective |
| Evidence horizon | months |
| Author role | engineering (internal-practice account: Moreno + Libbey) |
Theory coverage at a glance
Where the source’s described behavior maps onto each MAGE construct: ✓ strong · ◐ partial · ~ tension · ✗ counterexample · — not described.
| MAGE construct | Correspondence |
|---|---|
| Alignment | ✓ strong |
| Modeling | ◐ partial |
| Knowledge rep | ◐ partial |
| Bootstrap | ✓ strong |
| Conversion | ✓ strong |
| Determinization | ✓ strong |
| Reasoning horizon | ✓ strong |
| Engineer's seat | ✓ strong |
| Graduated governance | ◐ partial |