Industry case studies
MAGE did not run at this company. This page reads an independent practitioner report through MAGE’s vocabulary, to see how cleanly an outside system maps onto the theory. Every correspondence below is the book’s reading of the source, never a claim the company makes about MAGE.
These cases are practitioner reports, not replications of MAGE and not causal tests. A correspondence cell states how cleanly the source's described behavior maps onto a MAGE construct as the BOOK reads it — never a claim Cloudflare (or any site) makes about MAGE, and never proof. not-described means the SOURCE is silent; absence from a report is NOT evidence the company lacks the practice.
Meet the case
Zenseact carries MAGE past source code. Its agents do not write software; they do knowledge work across the company's own systems, reasoning over issue trackers, code review, CI, search, and even ERP to decide which tool and which piece of expertise a task needs. The platform, called ZAP, is federated: a central team owns the runtime, permissions, and context machinery, while each domain team ships an agent as little more than a directory holding a config file and a skill. The origin is telling: an earlier, open approach 'did not scale,' so Zenseact rebuilt the environment around the same commodity models, betting that capability lives partly in the environment, on a single company-authored source that reports no outcome numbers yet.
Distinctive starting point: autonomous knowledge work over distributed engineering systems
What the case shows
Zenseact built a federated internal agent platform - deterministic tools behind structured seams, a per-tool/per-role permission table, bounded delegation, and a deterministic TF-IDF disclosure router with measured 67-98% / 43-97% context reduction - over distributed engineering-and-enterprise knowledge work rather than source code, carrying MAGE's Alignment, reasoning-horizon, capital, and conversion machinery cleanly beyond code (the governance half of the MBAE widening, complementary to Siemens's modeling half), with the honest bound that its churn/defect outcome vocabulary strains on non-artifact-production work and no effectiveness outcomes are reported.
What the agents do
ZAP agents are more than document-retrieval chatbots: they use tools to act over enterprise systems and compose multi-step activity dynamically - deciding what information a task needs, which skill/tool is relevant, which action to perform, and whether to delegate to a specialist agent. Work leans toward engineering + enterprise knowledge work (investigation, synthesis, tool-mediated action across issue-tracking, code review, CI/CD, search, and financial/ERP systems), not autonomous source-code authorship; the autonomy is in reasoning about and orchestrating heterogeneous work.
Setting: org-wide · brownfield
Scale, as the source reports it:
- measured context reduction of 67-98% for financial queries and 43-97% for engineering queries via progressive skill disclosure (context cost measured; the reasoning-quality half asserted, not measured)
- skill frontmatter budget ~1.5-2 KB total across all active skills (always in the system prompt); skill bodies injected only when the deterministic TF-IDF router scores them relevant
- integrations span JIRA, Gerrit, Zuul, GoCD, Elasticsearch, SAP (ERP), Confluence, SharePoint, and proprietary internal tools, built on AWS Bedrock
- creating a domain agent requires only a directory under agents/ with two files (config.yaml + SKILL.md); no productivity/defect/before-after outcome data and no stated deployment duration
The engineered environment
Object territory
source-code design-docs incident-reports
Representations
informal-knowledge spec structured-policy
Mechanisms observed
retrieval-layer constrained-api merge-gate provenance stable-identity deterministic-lint
Where authority sits
Authority is explicit, machine-readable, and placed in the environment. Three stacked surfaces: a per-tool, per-role permission table (auto_approve / confirm-via-SSE / deny), so decide-freely != act-freely and the environment sets the line; agent-to-agent delegation that is allow-listed in config and depth-limited (the delegation topology is itself an authority model); and sanctioned tool seams with bidirectional sanitization (arguments validated for injection patterns; tool output scrubbed of prompt-injection markers). Boundaries live outside the reasoner.
Mapping into MAGE
Each row is one MAGE construct, the strength of the correspondence, and how the book reads the source against it. The note is the book’s reading; the strength is not a score.
| MAGE construct | Correspondence | How the book reads the case |
|---|---|---|
| Alignment | ✓ strong | the book reads the per-tool/per-role auto/confirm/deny table, allow-listed + depth-limited delegation, bidirectional sanitization, and deterministic tools behind structured seams as authority placed outside the reasoner - with Docker the set's sharpest Alignment instance, adding a graduated per-action ladder and delegation-topology-as-authority |
| Modeling | ◐ partial | the book reads rich structured knowledge representation (skills) but no invariant-bearing system/process models and no model<->territory join - skills describe how to work in a domain, not what the system is or must remain; a mild Modeling tension held at partial for matrix parity |
| Bootstrap | ✓ strong | the book reads the existing engineering + enterprise tool estate, team knowledge, and explicit policies as seeding E(0), with the failed open-source attempt as a bootstrap-by-failure origin |
| Conversion | ✓ strong | the book reads the disclosure-miss loop (miss -> telemetry -> automated skill-improvement pipeline) as a clean failure->diagnosis->environment-change->better-behavior sequence, narrowed to the routing/representation layer (the subtractive reconcile/retire arm is not-described) |
| Determinization | ✓ strong | the book reads three judgments other sites leave probabilistic being hardened into repeatable mechanisms - execution (deterministic tools), authority (the permission table), and context selection (the TF-IDF router) - a clean contrast with Docker |
| Reasoning horizon | ✓ strong | the book reads progressive disclosure with the set's strongest measured context reduction (67-98% / 43-97%) as reasoning-horizon management, noting the reasoning-quality half is asserted, not measured |
| Engineer's seat | ✓ strong | the book reads humans retaining skill/tool/permission authorship (distributed to domain teams) plus the confirm gate on consequential actions as the engineer's seat |
| Graduated governance | ◐ partial | the book reads auto/confirm/deny as a strong graduated-authority primitive across actions, but with no soft->hard mechanism-strengthening lifecycle over time (graduation across actions, not across maturity - like Docker) |
The theory the case appears to hold
The book reads Zenseact as holding that enterprise autonomy scales when the organization separates a common agent platform from domain-specific intelligence: teams encode expertise as reusable skills + deterministic tools; agents receive only task-relevant context rather than the whole enterprise per session; probabilistic models decide what knowledge/operation is needed, but consequential actions execute through structured tools under explicit permission policies; agent-to-agent delegation is bounded like tool access; and the platform is instrumented so context-disclosure and capability-selection failures become evidence for improving the shared environment.
What the case adds to MAGE
MAGE compresses Zenseact's separately-named skills / tools / agents / progressive disclosure / permissions / delegation / sanitization / monitoring into one governed environment - skills as soft conditioning, the TF-IDF router as a sensor + reasoning-horizon mechanism (and a determinization of context selection), tool schemas as sanctioned seams, the permission table as authority allocation + a human gate, the delegation allow-list as a constraint on the agent topology, sanitization as a trust boundary, disclosure-miss telemetry as a sensor, the improvement pipeline as governance conversion, and the platform/domain split as an agent stack over federated engineering capital - and names the structural-diagnosis question Zenseact performs but never states.
capability-discovery-itself-must-be-governed delegation-topology-is-an-authority-surface graduated-per-action-authority-auto-confirm-deny distributed-federated-governance-at-org-scale environment-improvement-can-target-context-selection trust-boundaries-run-in-both-directions the-context-router-can-be-deterministic
What MAGE adds that this case does not reach
MAGE is a theory of the engineering ENVIRONMENT ITSELF as the object of engineering: everyone else engineers an agent, a runtime, a policy engine, or a model; MAGE engineers the governed environment in which commodity intelligence operates.
None of the external cases we examined describes the following machinery in the generalized form MAGE does. Not 'nobody in industry has ever done this.'
Honest bounds
The limitations the analysis records, and the falsifiable hypotheses the case bears on:
no-invariant-bearing-system-model-described-modeling-tension no-model-territory-drift-gate churn-defect-outcome-vocabulary-strains-on-knowledge-work no-productivity-or-defect-or-before-after-outcome-data deployment-duration-unstated reasoning-quality-half-of-progressive-disclosure-unmeasured single-primary-source reconcile-retire-not-described
H2-governance-moderation H3-mechanized-assurance H4-representation-leverage H4a-representation-efficiency H6-oversight-amortization H7-conversion-conditions H8-learning-propagation
Source
| Field | Value |
|---|---|
| Citation | dufwa2026zap |
| Source type | build-report |
| Independence | independent |
| Account type | architectural |
| Evidence horizon | months |
| Author role | engineering (company-authored internal-platform whitepaper: Dufwa + Luvo) |
Theory coverage at a glance
Where the source’s described behavior maps onto each MAGE construct: ✓ strong · ◐ partial · ~ tension · ✗ counterexample · — not described.
| MAGE construct | Correspondence |
|---|---|
| Alignment | ✓ strong |
| Modeling | ◐ partial |
| Knowledge rep | — not-described |
| Bootstrap | ✓ strong |
| Conversion | ✓ strong |
| Determinization | ✓ strong |
| Reasoning horizon | ✓ strong |
| Engineer's seat | ✓ strong |
| Graduated governance | ◐ partial |