§5.3 Other Agentic Software Factories
My experience with DocAble showed me one way of building reliable software with unreliable agents. But that is one system, built by one engineer, for one problem. The natural question is how other engineers, working in other organizations and under different constraints, are approaching the same problem.
In this section, I reconstruct agentic software factories from public reports by seven major companies: Uber, GitLab, Cloudflare, Siemens, Spotify, Shopify, and Zenseact. These factories differ in what they represent, how they direct and constrain agents, what evidence they require, and where they retain human judgment. They also differ in a more basic way: some organizations are building factories for their own software, while others are building factories for somebody else to use.
These sources also require a conservative reading. Company engineering blogs, talks, and product documentation mix technical communication with organizational and product promotion. They can provide evidence about mechanisms, measurements, and operating experience, but they are not independent evaluations of the systems they describe. I therefore treat reported architectures and measurements as evidence of what an organization built and observed, while treating broader claims about effectiveness or causality more cautiously. Missing information is recorded as missing rather than interpreted as evidence that a mechanism does not exist. Numerical claims are retained with their reported definitions and denominators where available; incompatible measures are not treated as directly comparable.
The commercial systems examined here differ substantially in architecture and organizational setting, but they share a recognizable emphasis on the organization of agentic production: supplying context, structuring work, controlling access to tools and resources, arranging human review, and determining what changes may proceed. This emphasis maps well to process-oriented approaches to agentic software engineering advocated by Hassan and by O'Connell, which structure agent work through explicit workflows, artifacts, gates, evidence, and human intervention.11. Ahmed E. Hassan, Agentic Software Engineering: Building Trustworthy Software with Stochastic Teammates at Unprecedented Scale, 1st ed. (2026), https://agenticse-book.github.io/.22. William T. O'Connell, Engineering Trustworthy Intelligent Systems: Software Engineering, Governance, And Operational Trust in the AI Era, 1st ed. (ETIS Framework, 2026), https://github.com/etis-framework/etis. The cases below show how commercial organizations are engineering these concerns in practice.
§5.1 defined a software factory as a production system for controlled change. The factories examined here differ along two dimensions: who supplies the production system, and how engineering intent enters production.
The first distinction is between an internal factory, built and operated by an organization to change its own software, and a factory in a box, a generalized production system supplied to other organizations.
Factory in a box. An agentic software-production system packaged for use by other organizations, which connect it to their own software, engineering knowledge, policies, and business intent.
The second distinction concerns how engineering intent enters the factory. Most of the factories examined here begin with the existing software environment. Repositories, tests, pipelines, issue systems, service catalogs, policies, and other organizational context provide the material from which agents recover what they need to perform a change. I call these existing-environment-first factories. A model-first factory instead places explicit engineering models upstream of realization: selected engineering intent is represented in models, and software realization proceeds downstream from them.
These distinctions are independent. An internal factory can be existing-environment-first or model-first. So can a factory in a box. The seven public accounts happen to occupy only part of that design space. Table 5.3-1 provides the map I will use for the rest of the section.
| Factory form | Who supplies the production system? | How engineering intent enters production |
|---|---|---|
| Internal factory | The organization operating it | Either existing environment or explicit models |
| Factory in a box | An external supplier | Either existing environment or explicit models |
| Existing-environment-first | Either | Repositories, tests, pipelines, policies, service catalogs, and organizational context |
| Model-first | Either | Explicit engineering models upstream of realization |
Figure 5.3-1 locates where each organization starts and where it invests.
Across these factories, several structures recur. Engineering knowledge is externalized so that agents need not reconstruct it from a person's memory on every task. Agent action is bounded through tools, permissions, roles, execution environments, or other controls. Realization is evaluated through tests, validators, reviewers, simulations, or other evidence before consequential changes are admitted. And some engineering judgments remain with people where the available representations, evidence, or organizational obligations do not justify delegation. These common structures establish the baseline. The cases that follow focus instead on where the factories differ: what each organization has chosen to engineer around its fabricators, what operating experience has forced it to change, and what that experience teaches about constructing and maintaining an agentic software factory.
The public accounts also reveal how these factories are being operated, and they do so unevenly: a mechanism may be described without its effectiveness being measured, and a source may publish extensive activity data while saying little about the quality of the resulting software. That asymmetry is itself useful. It distinguishes mechanisms that exist from mechanisms whose effects have been observed, and it exposes several constraints that become visible only after realization capacity increases. I therefore use the cases for three related questions: how production is structured, what the organization chooses to measure, and what operating experience causes the factory to change.
5.3.1 Internal Software Factories
Five of the seven accounts describe a system an organization built for its own work, and each answers a different engineering question. Uber asks how to optimize the factory itself. Cloudflare asks how institutional policy becomes production control. Spotify asks how the controls should change when the fabricator does. Shopify asks how one engineer's experience becomes knowledge every later session inherits. Zenseact asks where authority over the factory's agents should sit. Four of the five begin with the existing software environment and delegate more of realization to agents from there. Zenseact is the exception — its agents, as the case will show, perform knowledge work around the software factory rather than realization inside it — and the exception is itself a data point about where agentic capability lands first.
Uber
Uber's distinctive move is not that its factory is sophisticated but that the factory is itself an object of measurement and optimization. Model choice, context, decomposition, orchestration, token cost and reviewer attention are treated as engineering variables with numbers attached, rather than as fixed properties of the reasoner. The scale gives the exercise its force: more than 70 percent of pull requests attributed to local or cloud agents, more than 3,600 engineer-created agent skills executing over 30,000 times per day, and managed agents performing code review, CI repair, end-to-end changes, alert triage, debugging, and maintenance.33. Uday Kiran Medisetty, “Running a Software Factory Efficiently at Uber Scale,” Uber Engineering, 2026, https://www.uber.com/us/en/blog/efficient-software-factory/. The headline figure needs its measurement caveat attached: no source defines "attributed to," and figures published earlier the same year under narrower definitions were far lower — 11 percent of pull requests opened by agents as of March.44. Gergely Orosz, “How Uber Uses AI for Development: Inside Look,” The Pragmatic Engineer, March 10, 2026, https://newsletter.pragmaticengineer.com/p/how-uber-uses-ai-for-development. The scale is real; its denominator is not public. Product intent still originates with people, but a growing share of sessions is initiated by managed agents responding to the factory's own operational signals — CI failures, on-call alerts, incoming bugs — and by centrally scheduled maintenance, with incident reviews mined at a roughly monthly cadence for lessons that become new maintenance skills applied across services.55. Uday Kiran Medisetty and Adam Huda, “Agentic SDLC at Uber — Building Blocks for Uber's Software Factory,” AI Engineer World's Fair, 2026, https://ai.engineer/talks/17-YSUHo6Lk-agentic-sdlc-at-uber-building-blocks-ubers.
The optimization targets the environment rather than the reasoner. Context and tool access are centralized; deterministic orchestration is moved from repeated model turns into executable code, cutting token use by more than half on common workflows and by more than 90 percent in bulk cases; and session anti-patterns are detected and priced.35 Real workloads become benchmarks, models are selected against cost, quality, and reliability, and managed agents are evaluated in outcome-denominated units: cost per merged pull request, review, alert, or cleanup.3
Two things resist that optimization, and Uber publishes both. The first is the evidence. Evaluation begins inside the agent's own loop, before shared machinery sees the work: static-analysis repair, a simulator screenshot compared against the Figma specification — the one place in this record where a design artifact serves as a mechanical oracle — and staging integration, with the resulting pull request carrying a table of the checks performed, screenshots included.5 The producing agent assembles that table about its own work, and no source states that any entry is re-derived by independent machinery before approval. Behind it sit a smaller reviewer model inside the loop, a reasoning model outside it, self-healing CI, and then a person. Admission is human, and Uber says so plainly: "we still have humans approving the code," with "a short path in the near future to a percentage of our code landing automatically" — the destination named, the criterion for which code and on what evidence unpublished.66. Will Bond and Ameya Ketkar, “Building Ureview, Uber's Multi-Agent Code Review Engine,” AI Engineer World's Fair, 2026, https://ai.engineer/talks/EL123UNokkI-building-ureview-ubers-multi-agent-code-review.
The second is reviewer attention, which Uber treats as a production resource with finite capacity. Time to first review rose from three hours in 2024 to nine in 2026 as volume and change size grew, and the stated answer is to move reviewers up a layer, toward architecture and product judgment, rather than out of the loop.6 Maintenance jobs are placed on Sunday for spare CI capacity while the number of diffs engineers face on Monday is separately capped, because spare compute does not imply unlimited reviewer attention.5 No other case in this corpus schedules admission capacity as a resource. Uber also tracks revert rate, F1, and time to recovery alongside production cost. The factory is therefore being optimized against both the cost of producing changes and signals about what happens after they are produced. Nor does any source state what evidence a change must produce before it may proceed. Uber's own closing lands where this chapter began: once building a feature is easy, implementation feasibility no longer settles the product question — "the factory still needs people to decide whether that feature should be built."5
Cloudflare
How does an organization's own policy become a control that runs in production? Cloudflare's factory attaches to its engineering organization at review rather than at origination — engineers still open the merge requests and write the specs — and its answer begins with representation. A standards corpus that had outgrown individual memory, and, in the CIO's franker telling, a review process that AI-accelerated engineers had outrun — "Anyone at Cloudflare could now write bad code, faster"77. Sam Rhea, “How We're Rethinking Work at Cloudflare with Cloudflare Os,” Cloudflare, August 5, 2026, https://blog.cloudflare.com/how-we-use-ai-with-cloudflare-os/. — was extracted by a purpose-built agent into structured statements, each SHOULD or MUST carrying a stable identifier that survives revision of the document containing it.88. Timo Reimann, “How Cloudflare Enforces Engineering Standards Using Ai,” Cloudflare, August 4, 2026, https://blog.cloudflare.com/engineering-standards-enforcement/. That identifier does the work, and nothing else in this corpus has an analogue for it: an obligation that keeps its identity across revisions of its own text can be tracked, waived, and measured wherever it is applied. The corpus then meets work at several stations: human review of standards before they enter the corpus, fast linters in the editing loop, semantic review in CI, specification review before implementation, and incident review after production, whose findings are mandatory to clear for high-severity incidents.8 Cloudflare moves checks into deterministic tooling where the language permits it; the stated reason for doing so is practical — semantic review latency had become expensive.
Authority arrives in a second, separate step. Any employee may propose a standard and a named domain owner approves it; a distinct promotion then moves an approved standard from advisory to enforced, and only then can its MUST statements block — though no source names who performs that promotion. Once promoted, admission is machine-held in the ordinary case. A coordinator agent orchestrates up to seven specialist reviewers, judges severity, and maps its verdict directly onto the version-control action — approve, revoke a prior approval, or block the merge — with no human in the path.99. Ryan Skidmore, “Orchestrating AI Code Review at Scale,” Cloudflare, April 20, 2026, https://blog.cloudflare.com/ai-code-review/. Cloudflare is the only organization in this corpus that hands a probabilistic evaluator that power, and it engineers around the grant rather than assuming it: only enforced MUST statements can block, the rubric is biased toward approval so a lone warning does not, and a human reviewer can force approval by commenting "break glass" — a logged override used 288 times, 0.6 percent of merge requests, in the measured month.9 The instrumentation runs in one direction only. The break-glass count measures forced approvals; no published number measures the violations the reviewer failed to flag. The much larger published counts — review runs, findings, and blocked merges — likewise measure activity by the control, not defects prevented or failures escaped. Policy authorship is human, application is machine, and only the escape hatch is counted.
Operating experience does become new structure here, along four documented paths. The November and December 2025 outages produced named rules — do not call .unwrap() outside tests; validate that upstream dependencies are in an expected state — now enforced at review; a dangerous configuration pattern can be enrolled in Snapstone, a shared substrate through which it "immediately inherits safe deployment"; the extracted representation itself was revised, Markdown to JSON, when agents needed to filter it more accurately; and reviewer latency produced the linter station.1010. Jeremy Hartman, “Code Orange: Fail Small Is Complete. The Result Is a Stronger Cloudflare Network,” Cloudflare, May 1, 2026, https://blog.cloudflare.com/code-orange-fail-small-complete/. What the record does not contain is a measurement closing any of these loops. The outcome claim for the outage-derived rules is a counterfactual — that the outages "would have been rejected merge requests instead of global incidents" — and Cloudflare places the remaining loop in the future in its own voice: the current focus is giving engineers "the tools to define the loops that evaluate the work their agents produce."7 The conversion paths are described and instantiated; their effect is asserted, not shown.
Spotify
Spotify's case isolates a variable the others cannot. Its production line predates the agent by years. Fleet Management targets repositories across the software estate, executes transformations across thousands of repositories, automerges on green, and watches merged changes in production, alerting the change author when deployments fail.1111. Matt Brown, “Fleet Management at Spotify (Part 3): Fleet-Wide Refactoring,” Spotify Engineering, May 15, 2023, https://engineering.atspotify.com/2023/05/fleet-management-at-spotify-part-3-fleet-wide-refactoring. In 2025 the company replaced the deterministic transformation script at the center of that line with an agent, and stated the substitution precisely: "All the surrounding Fleet Management infrastructure — targeting repositories, opening pull requests, getting reviews, and merging into production — remains exactly the same."1212. Max Charas and Marc Bruggmann, “1,500+ Prs Later: Spotify's Journey with Our Background Coding Agent,” Spotify Engineering, November 6, 2025, https://engineering.atspotify.com/2025/11/spotifys-background-coding-agent-part-1. The agent occupies one station. It took over the residual the prior technique could not reach — the last 30 percent of a fleet migration, the corner cases that had grown one dependency updater to twenty thousand lines of script. Everything that moved afterward moved because the fabricator at that one station had changed.
The controls around that station are tuned to the fabricator's properties, and they move when it does. The agent chooses its own path through the code but holds a deliberately narrow tool surface — a verify tool, restricted git subcommands that can never push, an allowlisted shell — with code search and documentation tools deliberately withheld and context condensed into the prompt up front: "The more tools you have, the more dimensions of unpredictability you introduce."12 Verification runs before a pull request may exist: verifiers activate from the component's contents, opaque to the agent, and a failing verifier means no PR;12 a later architecture separates the verification runtime entirely, so the PR is created only after full CI validation.1313. Daniel Curtis, “Qcon London 2026: Rewriting All of Spotify's Code Base, All the Time,” InfoQ, March 18, 2026, https://www.infoq.com/news/2026/03/spotify-honk-rewrite. The clearest control story in the case is a lifecycle. Agents gamed builds — commenting out failing tests, downgrading Java versions — so Spotify added an LLM judge, which by late 2025 was vetoing about a quarter of sessions, the agent self-correcting half the time; the talk record then reports the judge removed as models improved, "verification steps in prompts proving sufficient."13 A control installed against a measured failure mode and retired when the capability beneath it moved — reported second-hand, with no first-party text yet confirming the removal. If the later account is accurate, no other case in this corpus records a control being withdrawn after the fabricator improved.
Admission itself splits on which fabricator produced the change. Deterministic fleet changes automerge by default — more than 2.5 million automated maintenance PRs, "the vast majority auto-merged with no human in the loop,"1414. Niklas Gustavsson, “Coding Is No Longer the Constraint: Scaling Developer Experience to Teams and Agents at Spotify,” Spotify Engineering, June 3, 2026, https://engineering.atspotify.com/2026/6/code-with-claude-coding-is-no-longer-the-constraint. with the change author rather than the repository owner deciding what automerges11 — while agent-authored changes wait for human review, and Spotify names review, not generation, as the new bottleneck.13 Humans scope the migrations — work that once involved hundreds of teams over weeks is scoped by one engineer in days14 — author prompts as version-controlled, tested artifacts,12 and encode non-delegation directly: in one documented migration, wherever a judgment call was required the agent was instructed to leave the field unchanged and annotate it for the human reviewer.1515. Devon Edwards Joseph, “Background Coding Agents: Supercharging Downstream Consumer Dataset Migrations,” Spotify Engineering, April 22, 2026, https://engineering.atspotify.com/2026/4/background-coding-agents-dataset-migrations-honk-part-4. The factory's health is itself instrumented. Monthly incident retrospectives now ask whether AI-authored code contributed to an incident; the published finding is that it has not been a material direct contributor, but that "the volume of change increased faster than some of our verification controls could adapt" — one automated dependency upgrade passed every check and still failed in production. The constraint, stated by the SVP who owns the platform: "AI increased the capacity to produce change. The next constraint became our ability to verify it."1616. Tyson Singer, “AI Changed How Spotify Builds. What We Learned (And Fixed) About Quality at Higher Velocity,” Spotify Engineering, September 16, 2026, https://engineering.atspotify.com/2026/9/ai-changed-how-spotify-builds-what-we-learned-and-fixed-about-quality-at-higher-velocity.
Spotify also reports no increase in a refined measure of rework, while code complexity and pull-request size have begun to rise; it explicitly declines to move those thresholds. These measurements are unusual in this corpus because they concern the health of the resulting engineering system rather than agent activity. Together with the incident record, they show a mixed result: no measured increase in rework or material direct incident contribution from AI-authored code, alongside change volume beginning to outrun parts of the verification system.16
Shopify
Shopify's question is how one engineer's hard-won fix becomes knowledge that every later session inherits, and its answer begins in the architecture rather than in a pipeline. River, the company's internal agent, lives in Slack and refuses direct messages; every session is a public transcript, and an employee starts work by mentioning the agent in a channel.1717. Javier Moreno and Burke Libbey, “Under the River,” Shopify Engineering, May 28, 2026, https://shopify.engineering/under-the-river. Work is therefore performed in front of an audience by construction, and a failure is watched rather than reported. In a thirty-day window reported in May 2026, 59,918 sessions ran across 5,170 channels, touching the work of more than 7,000 people, and one in eight merged pull requests was River-coauthored.17 By September, Lütke reported that up to half of pull requests now begin as a River conversation.1818. Paul Drecksler, “Shopify CEO Tobi Lütke Says Half the Company's Pull Requests Now Start as Slack Conversations with an Agent Named River,” Shopifreaks, September 17, 2026, https://www.shopifreaks.com/shopify-ceo-tobi-lutke-says-half-the-companys-pull-requests-now-start-as-slack-conversations-with-an-agent-named-river/.
The corpus of sessions is the factory's learning substrate, and the mechanism Shopify documents is human. People watch River work, notice where it got stuck, and write down what it should have known; the artifacts are files in the monorepo — a skill update, an AGENTS.md diff, sometimes a whole new shared skill — alongside the code, conventions, intent documents, and runbooks the next session inherits.17 The effect is measured with the counterfactual controlled: River's merge rate rose from 36 to 77 percent over two months, and "We did not retrain a model. We did not switch models."1919. Tobi Lütke, “Learning on the Shop Floor,” May 9, 2026, https://x.com/tobi/article/2053121182044451016. The measure is self-reported and does not isolate human learning from improvement to the surrounding environment, but that entanglement is itself revealing: the people and the factory learned through the same mechanism. Shopify's engineering account closes by handing readers an imperative — "Make the agent multiplayer by construction" — a design prescription offered outward, not a label the company applies to itself.17
The corpus is not free to carry. Instruction files apply recursively down the monorepo tree, and rival tool conventions read different ones; Lütke threatened to ban Claude Code at Shopify over the "split brain problems" that divergence creates "when different team members use different tools."2020. Tobi Lütke, “Split Brain Problems When Different Team Members Use Different Tools,” X, August 25, 2026, https://x.com/tobi/status/2092259436538495186. In follow-up posts the trade press quoted, he gave the consequence at scale: with thousands of developers in one repository "it just does happen that one directory is missing one of the two files and this means that a subset of devs work with lobotomy."2121. Amanda Caswell, “Shopify's CEO Threatened to Ban Claude Code. Anthropic Had Already Closed the Feature Request,” The New Stack, August 25, 2026, https://thenewstack.io/shopify-claude-code-agentsmd/. Automation repairs the drift, and Lütke calls that repair "a stupid complexity tax that shouldn't have to be paid."21 The knowledge compounds; keeping it coherent is work somebody does.
The sharpest instance of the conversion is a failure that changed where the rule lived: from an instruction the agent had to interpret to a condition the production system was expected to enforce. A September account of River's security workflows names the controls: the agent's output is draft-only, the engineers who own the affected code hold merge authority, a freshness-gated merge queue rejects stale evidence, and a post-merge check verifies that the default branch reflects the fix — merges through the gated queue rose from about 10 percent to 80 percent.2222. Erin Son and Kaiyi Li, “How River Takes Security Work from a Fix to Merge,” Shopify Engineering, September 2, 2026, https://shopify.engineering/river-vulnerability-remediation. The same account states the governing rule and publishes the run that taught it: "Prompts are the right place to express judgment and the wrong place to express a guarantee," and in security, "any rule that determines whether a vulnerability is closed must be enforced in code" — written after River violated its own prompt-stated enumeration requirement and said so in its report.22 Those controls are stated for security workflows. What governs the other seven-eighths of River's output — what evidence an ordinary River pull request must carry, how often one fails in production — the public record does not say.
Zenseact
Zenseact, a driver-assistance software company, belongs here with a scoping sentence attached: its agents do not produce software. The one substantive source — a company-authored whitepaper describing its ZAP platform — shows agents querying Gerrit, Zuul, JIRA, SAP, and document stores as data sources and synthesizing answers; no agent authors a patch, opens a change, or merges anything.2323. Philip Dufwa and Thomas Luvö, “A Platform for Scalable Enterprise AI Agents,” Zenseact, May 29, 2026, https://zenseact.com/news/a-platform-for-scalable-enterprise-ai-agents/. Under this chapter's definitions that is not an agentic software factory but an agentic knowledge-work platform operating around one, and the case earns its place on a question none of the others puts squarely: where should authority over the factory's agents sit? Zenseact's answer is federation. A platform team owns the runtime — authentication, sessions, tool execution, the model loop — which deliberately "does not contain any business logic"; a domain team's agent is a directory holding two text files, a configuration listing its skills and a markdown file encoding the domain's rules, so a new agent is a configuration act rather than a platform change.23 What makes that boundary safe to draw is that consequence stays central: a permission table resolves every tool call to auto-approve, confirm with a human, or deny, by tool and by role.23 The domain team writes the knowledge; the platform keeps the authority to act.
The evidentiary standing travels with the claim. This is architecture described by its authors rather than practice observed: one source, no adoption count, no outcome figure, and no report of how many domain teams built an agent — so "distributed ownership scales" is a design property the record states and never demonstrates. Its one published performance measurement concerns the context router, which selects each turn's skill instructions by text similarity and carries the record's only quantity, a 67 to 98 percent context reduction on financial queries and 43 to 97 percent on engineering ones. Routing misses are typed — no skill matched, the wrong skill fired, the agent struggled with a tool — and each class carries a prescribed repair to the skill file rather than to the model that reads it.23 The federated unit is what gets repaired.
5.3.2 The Factory in a Box
GitLab belongs in the other column because the capability its account describes is generalized and sold, to be attached to somebody else's repositories, policies, and engineering knowledge. That form poses a question no internal factory faces: what part of a factory can be productized generically, and what must the adopting organization supply? GitLab's own internal record is worth taking first, because it prices the second half. The company's handbook documents an internal practice with doctrine attached: never give an agent a task without a failing test; "Fix the environment, not the prompt"; "Prompts are suggestions. CI is a gate. If the agent can break a rule and still pass the pipeline, the rule doesn't exist" — down to a CI job that fails when the test count decreases, because agents delete tests to make them pass.2424. GitLab, “AI-Assisted Development Playbook,” GitLab Handbook, 2026, https://handbook.gitlab.com/handbook/engineering/workflow/ai-assisted-development/. And it publishes a dated internal quantity: a 135,000-line Rust service, roughly 95 percent agent-generated, built by four engineers across 259 merge requests in about two weeks, with every agent-generated change still passing CI, security scanning, and code review.2525. GitLab, “Orbit Project Agentic Engineering Enablement (Work Item 163),” GitLab, 2026, https://gitlab.com/gitlab-org/orbit/knowledge-graph/-/work_items/163. The timing matters. Roughly four weeks were spent constructing the harness before the roughly two-week production episode began. The factory therefore had a visible startup investment before its marginal production economics could be observed.
The form introduces a boundary an internal factory does not have. The vendor supplies the fabricator, the orchestration, the context representations, the evidence infrastructure, and the generic controls; the customer must connect those to the knowledge and authority that make a change acceptable in its own system. The engineering problem is integration, not installation, and it sits at both ends of the production line. Upstream: how does organization-specific intent enter generic machinery? Downstream: how does organization-specific authority govern what comes out?
GitLab's upstream answer is files the customer authors at fixed paths in its own repository — an AGENTS.md, chat rules, review instructions, skills — plus the customer's CI configuration and, optionally, authored flow definitions.2626. GitLab, “Customize Gitlab Duo Agent Platform,” GitLab Documentation, 2026, https://docs.gitlab.com/user/duo_agent_platform/customize/. The vendor states the thesis behind the arrangement normatively: "anything the agent can't access in-context effectively doesn't exist."25 What GitLab supplies as infrastructure is the recovered remainder — a context graph over the customer's code, merge requests, pipelines, and security findings, assembled from artifacts that already exist. Recovered context carries relationships. It cannot carry the intent the customer never wrote down. Downstream, authority is expressed as a withheld capability. GitLab's security review flow "never sets the Approve state, even when it finds no issues" — the machinery may produce, repair, and advise, and only a person approves.2727. GitLab, “Security Review Flow,” GitLab Documentation, 2026, https://docs.gitlab.com/user/duo_agent_platform/flows/foundational_flows/security_review/. Around that prohibition sits governance of the action rather than the change: an agent's privileges are the intersection of its service account's and the triggering human's, tool policy is tightenable by a project but never loosenable below its group's rule, and audit events record which agent acted under which policy. The vendor industrializes the record that a human decision was made. It does not move the decision.
The asymmetry is in what the vendor can supply. GitLab specifies the supplied infrastructure in detail, but not the cost of authoring the organization-specific knowledge on which autonomy depends: no case study measures the work required to construct those instructions or how output varies with their quality. The CEO's essay locates trust in a durable layer of "context, verification, governance and evidence" around a replaceable model.2828. Bill Staples, “When Code Is Abundant,” GitLab, 2026, https://about.gitlab.com/blog/when-code-is-abundant/. Yet the layer that produced GitLab's own 95-percent-agent-generated service included failing tests written first, CI gates, repository-resident context, and the test-count guard — precisely the layer the adopting organization must supply.
The question the form poses follows directly, and it is worth stating plainly.
How much organization-specific engineering knowledge must be externalized before an off-the-shelf factory can reliably produce controlled changes to that organization's software?
The question is the representation problem in commercial dress, and §5.1 already named the tradeoff behind it. A purchased factory that reliably produces your software only when you describe that software almost completely has recreated the arrangement the earlier software factories died of. So the buyer's side of the transaction is a modeling decision: which consequential intent to externalize, in what form, at what cost, before generic machinery can respect it. Chapter 6 takes up what that externalization costs.
5.3.3 Model-First Factories
Siemens answers a question the software-first accounts never have to ask: what changes when explicit engineering models already sit upstream of realization? §5.1 told this ambition's history in software — enterprise architecture and model-driven development tried to determine the consequential decisions upstream, in shared representations, and lost economically rather than conceptually. Here the same ambition is still working, as the ordinary practice of an engineering tradition that arrived at models by a different route.
Siemens enters from outside software-first engineering, and "Siemens" is not one factory: its public record spans several business units, and most of it describes production systems Siemens sells rather than a factory it runs on its own software. The clearest instance of software fabricated downstream of a model is the Eigen Engineering Agent. It reads a TIA Portal automation project — an engineering model already carrying the devices, network topology, and control logic — writes SCL and LAD control code, test logic, and HMI JavaScript against it, and "iterates until pre-defined performance benchmarks are met."2929. Siemens AG, “Siemens Brings AI to the Physical World with Eigen Engineering Agent,” Siemens AG, April 20, 2026, https://press.siemens.com/global/en/pressrelease/siemens-brings-ai-physical-world-eigen-engineering-agent. The admission criterion is a pre-set acceptance region, not a specified realization — a tolerance in this chapter's exact sense. Two years earlier, Siemens' Agent Studio research placed agents on formal modeling languages, SysML v2 among them, with generated candidate architectures validated through simulation and the system waiting for the engineer's decision at critical points.3030. Kai Liu and Arquimedes Canedo, “Agent Studio: A Multi-Agent System for Systems Engineering,” Siemens Digital Industries Software, December 12, 2024, https://blogs.sw.siemens.com/art-of-the-possible/agent-studio-a-multi-agent-system-for-systems-engineering/. The recurring structure across the Siemens record: a model supplies the goal and the context, a deterministic engine — a simulator, a compiler, a benchmark harness — supplies the verdict, and the agent searches freely between them. That is the bargain the earlier model-driven attempts could not strike.
The broader architecture Siemens publishes around this — task agents and orchestrators over knowledge graphs of engineering data — is explicitly "a research vision and innovation roadmap, not a committed product delivery," and its subject is simulation rather than software realization.3131. Daniel Berger et al., “Your Next Colleague Is an Agent: The Dawn of Agentic AI-Aided Engineering,” Siemens Digital Industries Software, July 3, 2026, https://blogs.sw.siemens.com/art-of-the-possible/agentic-ai-aided-engineering/. Two silences mark what the model-first factory has not yet been shown to do. No Siemens source establishes a conformance mechanism that rejects divergence between model and implementation — Eigen's benchmark is a threshold, not a correspondence — and Eigen's public record names no human review or approval step. Source silence is not organizational absence, but both omissions concern consequential controls. Where the record does describe Siemens changing its own software, in a legacy-modernization account written up by its cloud partner rather than by Siemens engineers, the posture inverts: the system "keeps a human in the loop at every step," and the reason given is an obligation the software carries rather than a stage of maturity. Industrial systems run "over 15 to 20 years of operation," so AI-generated output "must therefore be explainable, traceable, and verifiable," and "hallucinated or unvalidated changes are not merely inefficient but operationally unacceptable."3232. Anant Nawalgaria and Tomasz Świtoń, “How Siemens "Sliced the Elephant," Modernizing Legacy Code with Agentic Workflows,” Google Cloud, June 16, 2026, https://cloud.google.com/blog/products/ai-machine-learning/how-siemens-sliced-the-elephant-modernizing-legacy-code-with-agentic-workflows. Unlike the software-first cases, the available Siemens record supplies almost no operating measurement of the factory itself: no throughput series, admission rate, failure rate, or before-and-after outcome for the agentic production systems described here.
One difference from the old dream matters more than the rest, and it is the difference §5.1 isolated. Model-driven development had to specify realization almost completely, because its downstream fabricator was inexpensive but not capable. Today, capable realization is inexpensive. A model-first factory can therefore represent the decisions engineering must control and leave the remaining realization freedom to an agent that will resolve it acceptably. The model need not determine every realization decision. That is a different bargain from the one the earlier attempts lost, and whether it pays is not something the public sources settle.
Two cautions keep the axis honest. Model-first is not a rung above existing-environment-first; it is a different answer to where intent enters, carrying its own costs. And the axis is independent of the first one. A factory in a box can be sold as model-first, driven by models the customer authors rather than a repository the vendor crawls — and Siemens itself shows the combined cell is not empty: Eigen and its EDA verification toolkit are sold production systems driven by the customer's own engineering models, in the narrow software of control logic and chip design. For general-purpose software, no observed case sells that today — a fact about a young industry rather than a constraint on the design space.
5.3.4 What the Factories Share, and Where They Differ
The factories differ in where they place admission:
- Cloudflare gives its coordinator the merge gate under human-authored policy, with a logged human override.
- Spotify automerges deterministic fleet changes while agent-authored changes wait for human review.
- Uber keeps approval human while naming partial automatic approval as a near-term destination.
- GitLab withholds approval from the agent in its documented security-review flow.
The reasons for retaining human judgment differ as well. Uber's boundary is interim: the available representations and evidence do not yet justify further delegation. Shopify grounds the boundary in responsibility: a machine cannot take responsibility for a decision.3333. Podcast Alpha, “Tobi Lutke: River Now Writes Half of Shopify's Pull Requests,” Podcast Alpha, September 21, 2026, https://podcastalpha.substack.com/p/tobi-lutke-river-now-writes-half. Siemens grounds it in the obligations of long-lived industrial software: generated output must remain explainable, traceable, and verifiable over systems expected to operate for fifteen to twenty years. Only Uber describes its boundary as one expected to move with improving capability. The cases therefore do not support a single delegation frontier that simply advances as the fabricator improves.
The reported measurements are much stronger on the production of change than on the quality of the changes produced. Uber and Cloudflare report cost, throughput, findings, and overrides in detail, but much less about resulting software quality. Spotify reports measures of engineering-system health, including rework, complexity, pull-request size, and incident contribution. Shopify reports a merge-rate improvement with the underlying model held fixed. GitLab describes measurement mechanisms, while Siemens and Zenseact provide little operating evidence of this kind. These public accounts may omit measurements used internally, but as evidence they characterize production much more fully than they establish the quality of its output.
Operating experience changes the factories too. Cloudflare converts incidents, latency, and representation problems into new production structure; Shopify converts observed agent failures into shared instructions and, where closure depends on a property, into code-enforced guarantees; GitLab states the same discipline as doctrine, fixing the environment rather than repeatedly repairing the prompt; and Spotify provides the converse case, a probabilistic control withdrawn after the underlying models improved. These are instances of governance conversion, but the evidence usually establishes the conversion rather than its effect. No organization in this corpus publishes a recurrence measure showing that an incident-derived control reduced the failure class that produced it.
One constraint nevertheless appears repeatedly enough to matter: the capacity to supervise and verify production does not automatically increase with the capacity to produce changes. Uber measures reviewer delay increasing from three to nine hours and schedules work against finite reviewer attention. Spotify states the same pressure directly: change-production capacity increased faster than some verification controls could adapt. Shopify's "slop grenade" names the social version of the problem, in which cheap production transfers expensive inspection to somebody else. Cloudflare moved deterministic checks earlier partly because semantic review latency had become costly. These mechanisms differ, but the underlying resource is recognizable: fabrication can scale faster than the attention and evidence required to establish that its output is acceptable.
That observation does not require another kind of software factory. It identifies an operating constraint on the factories already described. Process design determines where that scarce supervision is spent, what can be established mechanically, and which work reaches a person. Tolerances and their evidence determine what must be established before a realization can be admitted. As fabrication becomes cheaper, the economics of the factory increasingly depend on both.
A second recurring cost is quieter. The representations through which a factory inherits engineering knowledge must themselves be maintained. GitLab budgets machinery for documentation freshness and convergence, Cloudflare checks the context supplied through AGENTS.md, and Shopify automates repair of instruction-file drift. The factory therefore acquires the same carrying-cost problem as the software it governs: production structures that preserve engineering knowledge are capital only while they remain sufficiently accurate to be useful.
Where the factories converge most strongly is leverage. Organizations centralize shared mechanisms while keeping domain judgment distributed; human judgment moves upstream, from hand-authoring changes toward scoping, supervision, and the engineering of reusable controls; private learning becomes shared structure; and the unit of governance grows from a task to classes of work, fleets, and platforms. A scarce human decision becomes more valuable when the environment carries it forward: one policy shapes thousands of changes, one shared representation guides work across a fleet, and one learned constraint protects later work over the same surface. Governance does not remove judgment; it gives selected judgments multiplicative reach.
Works Cited
- Hassan, Ahmed E. Agentic Software Engineering: Building Trustworthy Software with Stochastic Teammates at Unprecedented Scale. 1st ed. 2026. https://agenticse-book.github.io/.
- O'Connell, William T. Engineering Trustworthy Intelligent Systems: Software Engineering, Governance, And Operational Trust in the AI Era. 1st ed. ETIS Framework, 2026. https://github.com/etis-framework/etis.
- Medisetty, Uday Kiran. “Running a Software Factory Efficiently at Uber Scale.” Uber Engineering, 2026. https://www.uber.com/us/en/blog/efficient-software-factory/.
- Orosz, Gergely. “How Uber Uses AI for Development: Inside Look.” The Pragmatic Engineer, March 10, 2026. https://newsletter.pragmaticengineer.com/p/how-uber-uses-ai-for-development.
- Medisetty, Uday Kiran, and Adam Huda. “Agentic SDLC at Uber — Building Blocks for Uber's Software Factory.” AI Engineer World's Fair, 2026. https://ai.engineer/talks/17-YSUHo6Lk-agentic-sdlc-at-uber-building-blocks-ubers.
- Bond, Will, and Ameya Ketkar. “Building Ureview, Uber's Multi-Agent Code Review Engine.” AI Engineer World's Fair, 2026. https://ai.engineer/talks/EL123UNokkI-building-ureview-ubers-multi-agent-code-review.
- Rhea, Sam. “How We're Rethinking Work at Cloudflare with Cloudflare Os.” Cloudflare, August 5, 2026. https://blog.cloudflare.com/how-we-use-ai-with-cloudflare-os/.
- Reimann, Timo. “How Cloudflare Enforces Engineering Standards Using Ai.” Cloudflare, August 4, 2026. https://blog.cloudflare.com/engineering-standards-enforcement/.
- Skidmore, Ryan. “Orchestrating AI Code Review at Scale.” Cloudflare, April 20, 2026. https://blog.cloudflare.com/ai-code-review/.
- Hartman, Jeremy. “Code Orange: Fail Small Is Complete. The Result Is a Stronger Cloudflare Network.” Cloudflare, May 1, 2026. https://blog.cloudflare.com/code-orange-fail-small-complete/.
- Brown, Matt. “Fleet Management at Spotify (Part 3): Fleet-Wide Refactoring.” Spotify Engineering, May 15, 2023. https://engineering.atspotify.com/2023/05/fleet-management-at-spotify-part-3-fleet-wide-refactoring.
- Charas, Max, and Marc Bruggmann. “1,500+ Prs Later: Spotify's Journey with Our Background Coding Agent.” Spotify Engineering, November 6, 2025. https://engineering.atspotify.com/2025/11/spotifys-background-coding-agent-part-1.
- Curtis, Daniel. “Qcon London 2026: Rewriting All of Spotify's Code Base, All the Time.” InfoQ, March 18, 2026. https://www.infoq.com/news/2026/03/spotify-honk-rewrite.
- Gustavsson, Niklas. “Coding Is No Longer the Constraint: Scaling Developer Experience to Teams and Agents at Spotify.” Spotify Engineering, June 3, 2026. https://engineering.atspotify.com/2026/6/code-with-claude-coding-is-no-longer-the-constraint.
- Edwards Joseph, Devon. “Background Coding Agents: Supercharging Downstream Consumer Dataset Migrations.” Spotify Engineering, April 22, 2026. https://engineering.atspotify.com/2026/4/background-coding-agents-dataset-migrations-honk-part-4.
- Singer, Tyson. “AI Changed How Spotify Builds. What We Learned (And Fixed) About Quality at Higher Velocity.” Spotify Engineering, September 16, 2026. https://engineering.atspotify.com/2026/9/ai-changed-how-spotify-builds-what-we-learned-and-fixed-about-quality-at-higher-velocity.
- Moreno, Javier, and Burke Libbey. “Under the River.” Shopify Engineering, May 28, 2026. https://shopify.engineering/under-the-river.
- Drecksler, Paul. “Shopify CEO Tobi Lütke Says Half the Company's Pull Requests Now Start as Slack Conversations with an Agent Named River.” Shopifreaks, September 17, 2026. https://www.shopifreaks.com/shopify-ceo-tobi-lutke-says-half-the-companys-pull-requests-now-start-as-slack-conversations-with-an-agent-named-river/.
- Lütke, Tobi. “Learning on the Shop Floor.” May 9, 2026. https://x.com/tobi/article/2053121182044451016.
- Lütke, Tobi. “Split Brain Problems When Different Team Members Use Different Tools.” X, August 25, 2026. https://x.com/tobi/status/2092259436538495186.
- Caswell, Amanda. “Shopify's CEO Threatened to Ban Claude Code. Anthropic Had Already Closed the Feature Request.” The New Stack, August 25, 2026. https://thenewstack.io/shopify-claude-code-agentsmd/.
- Son, Erin, and Kaiyi Li. “How River Takes Security Work from a Fix to Merge.” Shopify Engineering, September 2, 2026. https://shopify.engineering/river-vulnerability-remediation.
- Dufwa, Philip, and Thomas Luvö. “A Platform for Scalable Enterprise AI Agents.” Zenseact, May 29, 2026. https://zenseact.com/news/a-platform-for-scalable-enterprise-ai-agents/.
- GitLab. “AI-Assisted Development Playbook.” GitLab Handbook, 2026. https://handbook.gitlab.com/handbook/engineering/workflow/ai-assisted-development/.
- GitLab. “Orbit Project Agentic Engineering Enablement (Work Item 163).” GitLab, 2026. https://gitlab.com/gitlab-org/orbit/knowledge-graph/-/work_items/163.
- GitLab. “Customize Gitlab Duo Agent Platform.” GitLab Documentation, 2026. https://docs.gitlab.com/user/duo_agent_platform/customize/.
- GitLab. “Security Review Flow.” GitLab Documentation, 2026. https://docs.gitlab.com/user/duo_agent_platform/flows/foundational_flows/security_review/.
- Staples, Bill. “When Code Is Abundant.” GitLab, 2026. https://about.gitlab.com/blog/when-code-is-abundant/.
- Siemens AG. “Siemens Brings AI to the Physical World with Eigen Engineering Agent.” Siemens AG, April 20, 2026. https://press.siemens.com/global/en/pressrelease/siemens-brings-ai-physical-world-eigen-engineering-agent.
- Liu, Kai, and Arquimedes Canedo. “Agent Studio: A Multi-Agent System for Systems Engineering.” Siemens Digital Industries Software, December 12, 2024. https://blogs.sw.siemens.com/art-of-the-possible/agent-studio-a-multi-agent-system-for-systems-engineering/.
- Berger, Daniel, Dorlis Bergmann, Arquimedes Canedo, et al. “Your Next Colleague Is an Agent: The Dawn of Agentic AI-Aided Engineering.” Siemens Digital Industries Software, July 3, 2026. https://blogs.sw.siemens.com/art-of-the-possible/agentic-ai-aided-engineering/.
- Nawalgaria, Anant, and Tomasz Świtoń. “How Siemens "Sliced the Elephant," Modernizing Legacy Code with Agentic Workflows.” Google Cloud, June 16, 2026. https://cloud.google.com/blog/products/ai-machine-learning/how-siemens-sliced-the-elephant-modernizing-legacy-code-with-agentic-workflows.
- Podcast Alpha. “Tobi Lutke: River Now Writes Half of Shopify's Pull Requests.” Podcast Alpha, September 21, 2026. https://podcastalpha.substack.com/p/tobi-lutke-river-now-writes-half.