Skip to content
All guidesAI systems

Software factories, explained

A software factory is a set-up where AI agents write the code and a chain of checks decides what ships. Here are its three loops, what it is for, its pros and cons, and the checks it cannot do without.

7 min read

More software is now written by AI agents, programs that can plan a task, write code, run it and fix what breaks without a person typing each step. When agents do most of the writing, something other than the writer has to decide what is good enough to ship. A software factory is a set-up built around that decision.

This guide explains the three loops inside a software factory, what factories are designed to achieve, where they help and where they hurt, and the checks that make one safe to trust. We run one ourselves to build Pulsar's own software, so the views at the end come from running it.

What a software factory is

The term is older than AI. US defence teams such as the Air Force's Kessel Run used it from 2017 for automated pipelines that take code from a developer to a live system with as little manual work as possible. In 2026 it took on a sharper meaning. StrongDM published a description of its own software factory, in which written specifications and test scenarios drive agents that write the code, and humans neither write nor review it line by line.

Most teams sit somewhere in between. Agents write much of the code, people set the goals and approve risky changes, and automated checks do most of the judging. What makes it a factory is that the route from idea to live system is the same every time, and each step is done or checked by a machine.

The three loops

A loop here means a cycle that repeats until something passes. A factory runs three of them, each slower and wider than the last. The terms inner loop and outer loop are common industry usage, and Microsoft's documentation is a good plain reference for them.

How one change moves through a software factory
  1. The agent writes and tests (inner loop)Agent

    The agent writes code, runs the tests, reads the errors and tries again, often in seconds. It works in a sandbox, a sealed copy of the software where mistakes cannot reach customers.

    sandboxtest suite

  2. The agent opens a pull requestAgent

    A pull request is a proposed change, packaged so others can see exactly what would change before it is accepted.

    GitHub

  3. Automatic checks run (outer loop)Agent

    Continuous integration, or CI, runs the full set of tests and scans on every change, the same way every time.

    CI

  4. A different reviewer checks it (outer loop)Agent

    A reviewer that did not build the change checks it against what was asked. This is usually a second agent with different instructions, with a person added for risky changes.

  5. A person approves risky changesYou

    Changes that touch payments, security or customer data wait for a named person to approve them. The agent that built the change never approves it.

  6. The change joins the main code and goes liveAgent

    The change joins the main code and goes live, ideally to a small share of users first.

    hosting

  7. The live system reports back (monitoring)Agent

    Error alerts, slow pages, usage and customer reports are watched continuously. A real problem becomes a new ticket, a logged to-do item, and the inner loop starts again.

    monitoringticket tracker

Simplified. Real factories add security scans, staging copies and approval rules for sensitive changes.

The inner loop

This is where the agent does its work. It is fast and cheap, and it is also where the agent is marking its own homework. Passing its own tests only shows that the code matches what the agent thought was wanted. That is why nothing leaves the inner loop on the agent's word alone.

The outer loop

This is the team's loop, and it takes hours rather than seconds. Its job is to check the change from outside. Automatic checks catch what can be tested mechanically. The independent review catches what tests miss, such as a change that passes every test but solves the wrong problem.

Monitoring

Some defects will always reach real users. Google's guidance on site reliability says as much, which is why it recommends releasing changes to a small group first. Monitoring closes the loop. When something goes wrong in the live system, the factory turns it into the next job rather than waiting for a customer to complain. DORA, Google's long-running research programme on software delivery, measures this with five numbers, including how often a change fails and how long recovery takes.

What software factories are designed to achieve

  • Small changes, shipped often. Small changes are easier to check and easier to undo.
  • The same checks on every change. A busy afternoon does not mean a skipped test.
  • More output from a small team. People spend their time on goals and judgement rather than typing.
  • A full record. Every change arrives with its tests, its review and its reason attached.

These are design goals. We have not found a reliable public study that measures, for example, how much a factory lowers the cost of each change, so treat any precise saving you hear with care.

Pros and cons

The pros follow from the goals. Work moves from idea to live system faster, standards are applied the same way every time, and there is an audit trail for everything. For a small business, this can put the capability of a much larger development team within reach.

The cons are just as real.

  • Mistakes ship faster too. A factory without strong checks spreads bugs at the same speed as features. DORA's 2025 report still found AI adoption linked to less stable software delivery, and summed it up by saying AI amplifies what is already there.
  • Agents can game tests. Research from METR and Anthropic has documented models editing or bypassing tests so that they appear to pass. A test the agent can see and change is a weak test.
  • AI-written code carries security risk. In Veracode's 2025 tests across more than 100 models, 45% of code samples failed security checks.
  • People stop reading the code. When review feels optional, understanding of the system fades, and the first hard bug takes much longer to fix.
  • It costs money to run. Agents are billed by usage. StrongDM suggests a mature factory spends at least $1,000 a day on AI usage per engineer, a figure the developer and writer Simon Willison has openly questioned.

The verification a factory needs

Sonar's January 2026 survey of more than 1,100 developers found that 96% do not fully trust AI-written code, yet only 48% always check it before saving it into their projects. A factory closes that gap by making the checks structural rather than optional. These are the five we would never skip.

  1. Tests first. Write down what done means, as tests, before any code exists, so the agent is aiming at a fixed target.
  2. A different checker. The agent that builds a change never approves it. Use a separate reviewer with its own instructions.
  3. Hidden scenarios. Keep some tests where the building agent cannot see them. StrongDM stores its scenarios outside the codebase for this reason, much like a held-back exam paper.
  4. Small rollouts. Try the change on a test copy, then on a small share of real users, and roll back quickly if the numbers move the wrong way.
  5. A person on risky changes. Anything touching money, personal data, security or deleting records needs a human yes.

For keeping AI-written code in good shape over time, see [how to keep AI-generated code maintainable](/learn/how-to-keep-ai-generated-code-maintainable). If you are deciding how an agent should reach your tools in the first place, [MCP or CLI](/learn/mcp-vs-cli) covers that choice.