Industry case studies

ReconstructionZenseact · engineering knowledge work / automotive

MAGE did not run at this company. This page reads an independent practitioner report through MAGE’s vocabulary, to see how cleanly an outside system maps onto the theory. Every correspondence below is the book’s reading of the source, never a claim the company makes about MAGE.

These cases are practitioner reports, not replications of MAGE and not causal tests. A correspondence cell states how cleanly the source's described behavior maps onto a MAGE construct as the BOOK reads it — never a claim Cloudflare (or any site) makes about MAGE, and never proof. not-described means the SOURCE is silent; absence from a report is NOT evidence the company lacks the practice.

Meet the case

Zenseact carries MAGE past source code. Its agents do not write software; they do knowledge work across the company's own systems, reasoning over issue trackers, code review, CI, search, and even ERP to decide which tool and which piece of expertise a task needs. The platform, called ZAP, is federated: a central team owns the runtime, permissions, and context machinery, while each domain team ships an agent as little more than a directory holding a config file and a skill. The origin is telling: an earlier, open approach 'did not scale,' so Zenseact rebuilt the environment around the same commodity models, betting that capability lives partly in the environment, on a single company-authored source that reports no outcome numbers yet.

Distinctive starting point: autonomous knowledge work over distributed engineering systems

What the case shows

Zenseact built a federated internal agent platform - deterministic tools behind structured seams, a per-tool/per-role permission table, bounded delegation, and a deterministic TF-IDF disclosure router with measured 67-98% / 43-97% context reduction - over distributed engineering-and-enterprise knowledge work rather than source code, carrying MAGE's Alignment, reasoning-horizon, capital, and conversion machinery cleanly beyond code (the governance half of the MBAE widening, complementary to Siemens's modeling half), with the honest bound that its churn/defect outcome vocabulary strains on non-artifact-production work and no effectiveness outcomes are reported.

What the agents do

ZAP agents are more than document-retrieval chatbots: they use tools to act over enterprise systems and compose multi-step activity dynamically - deciding what information a task needs, which skill/tool is relevant, which action to perform, and whether to delegate to a specialist agent. Work leans toward engineering + enterprise knowledge work (investigation, synthesis, tool-mediated action across issue-tracking, code review, CI/CD, search, and financial/ERP systems), not autonomous source-code authorship; the autonomy is in reasoning about and orchestrating heterogeneous work.

Setting: org-wide · brownfield

Scale, as the source reports it:

The engineered environment

Object territory

source-code design-docs incident-reports

Representations

informal-knowledge spec structured-policy

Mechanisms observed

retrieval-layer constrained-api merge-gate provenance stable-identity deterministic-lint

Where authority sits

Authority is explicit, machine-readable, and placed in the environment. Three stacked surfaces: a per-tool, per-role permission table (auto_approve / confirm-via-SSE / deny), so decide-freely != act-freely and the environment sets the line; agent-to-agent delegation that is allow-listed in config and depth-limited (the delegation topology is itself an authority model); and sanctioned tool seams with bidirectional sanitization (arguments validated for injection patterns; tool output scrubbed of prompt-injection markers). Boundaries live outside the reasoner.

Mapping into MAGE

Each row is one MAGE construct, the strength of the correspondence, and how the book reads the source against it. The note is the book’s reading; the strength is not a score.

MAGE constructCorrespondenceHow the book reads the case
Alignment✓ strongthe book reads the per-tool/per-role auto/confirm/deny table, allow-listed + depth-limited delegation, bidirectional sanitization, and deterministic tools behind structured seams as authority placed outside the reasoner - with Docker the set's sharpest Alignment instance, adding a graduated per-action ladder and delegation-topology-as-authority
Modeling◐ partialthe book reads rich structured knowledge representation (skills) but no invariant-bearing system/process models and no model<->territory join - skills describe how to work in a domain, not what the system is or must remain; a mild Modeling tension held at partial for matrix parity
Bootstrap✓ strongthe book reads the existing engineering + enterprise tool estate, team knowledge, and explicit policies as seeding E(0), with the failed open-source attempt as a bootstrap-by-failure origin
Conversion✓ strongthe book reads the disclosure-miss loop (miss -> telemetry -> automated skill-improvement pipeline) as a clean failure->diagnosis->environment-change->better-behavior sequence, narrowed to the routing/representation layer (the subtractive reconcile/retire arm is not-described)
Determinization✓ strongthe book reads three judgments other sites leave probabilistic being hardened into repeatable mechanisms - execution (deterministic tools), authority (the permission table), and context selection (the TF-IDF router) - a clean contrast with Docker
Reasoning horizon✓ strongthe book reads progressive disclosure with the set's strongest measured context reduction (67-98% / 43-97%) as reasoning-horizon management, noting the reasoning-quality half is asserted, not measured
Engineer's seat✓ strongthe book reads humans retaining skill/tool/permission authorship (distributed to domain teams) plus the confirm gate on consequential actions as the engineer's seat
Graduated governance◐ partialthe book reads auto/confirm/deny as a strong graduated-authority primitive across actions, but with no soft->hard mechanism-strengthening lifecycle over time (graduation across actions, not across maturity - like Docker)

The theory the case appears to hold

The book reads Zenseact as holding that enterprise autonomy scales when the organization separates a common agent platform from domain-specific intelligence: teams encode expertise as reusable skills + deterministic tools; agents receive only task-relevant context rather than the whole enterprise per session; probabilistic models decide what knowledge/operation is needed, but consequential actions execute through structured tools under explicit permission policies; agent-to-agent delegation is bounded like tool access; and the platform is instrumented so context-disclosure and capability-selection failures become evidence for improving the shared environment.

What the case adds to MAGE

MAGE compresses Zenseact's separately-named skills / tools / agents / progressive disclosure / permissions / delegation / sanitization / monitoring into one governed environment - skills as soft conditioning, the TF-IDF router as a sensor + reasoning-horizon mechanism (and a determinization of context selection), tool schemas as sanctioned seams, the permission table as authority allocation + a human gate, the delegation allow-list as a constraint on the agent topology, sanitization as a trust boundary, disclosure-miss telemetry as a sensor, the improvement pipeline as governance conversion, and the platform/domain split as an agent stack over federated engineering capital - and names the structural-diagnosis question Zenseact performs but never states.

capability-discovery-itself-must-be-governed delegation-topology-is-an-authority-surface graduated-per-action-authority-auto-confirm-deny distributed-federated-governance-at-org-scale environment-improvement-can-target-context-selection trust-boundaries-run-in-both-directions the-context-router-can-be-deterministic

What MAGE adds that this case does not reach

MAGE is a theory of the engineering ENVIRONMENT ITSELF as the object of engineering: everyone else engineers an agent, a runtime, a policy engine, or a model; MAGE engineers the governed environment in which commodity intelligence operates.

None of the external cases we examined describes the following machinery in the generalized form MAGE does. Not 'nobody in industry has ever done this.'

Honest bounds

The limitations the analysis records, and the falsifiable hypotheses the case bears on:

no-invariant-bearing-system-model-described-modeling-tension no-model-territory-drift-gate churn-defect-outcome-vocabulary-strains-on-knowledge-work no-productivity-or-defect-or-before-after-outcome-data deployment-duration-unstated reasoning-quality-half-of-progressive-disclosure-unmeasured single-primary-source reconcile-retire-not-described

H2-governance-moderation H3-mechanized-assurance H4-representation-leverage H4a-representation-efficiency H6-oversight-amortization H7-conversion-conditions H8-learning-propagation

Source

FieldValue
Citationdufwa2026zap
Source typebuild-report
Independenceindependent
Account typearchitectural
Evidence horizonmonths
Author roleengineering (company-authored internal-platform whitepaper: Dufwa + Luvo)

Theory coverage at a glance

Where the source’s described behavior maps onto each MAGE construct: ✓ strong · ◐ partial · ~ tension · ✗ counterexample · — not described.

MAGE constructCorrespondence
Alignment✓ strong
Modeling◐ partial
Knowledge rep— not-described
Bootstrap✓ strong
Conversion✓ strong
Determinization✓ strong
Reasoning horizon✓ strong
Engineer's seat✓ strong
Graduated governance◐ partial