Yonatan Boguslavsky
Opinion

A room full of agents is not a software factory

"My advice," writes Yonatan Boguslavsky, co-founder and CTO of Port, "is don't start with a grand design. Find out where agents are waiting on a person, on missing context, or on manual work."

The problem with agents in the SDLC is not the agents. They're good enough. They write code, review it, triage tickets, patch CVEs, and most engineering teams already use a few of them. The problem is what happens between the agents, or the lack thereof. Like when an alert fires and someone has to copy the error into a prompt. Or a ticket that says "fix login bug" and a person has to work out which repo that means. An agent might open a PR and it will sit in a queue because nobody is sure who owns that service. The agents are ready to move fast. But everything around it isn't.
In other words, what we’ve done in the SDLC so far is the equivalent of attaching a V12 engine to a bicycle. It wants to move fast but doesn’t have the foundation to support it. A software factory is what fixes it in the SDLC, and since the term is getting thrown around loosely, here's what it actually means.
1 View gallery
Yonatan Boguslavsky
Yonatan Boguslavsky
Yonatan Boguslavsky
(Nir Slakman)
What a software factory is
A software factory is a system where work moves across the SDLC through repeatable agentic workflows, with context, evaluation, and control built into the system itself.
An efficient version of it would look like this. A ticket arrives, the system works out which service it touches and who owns it, a coding agent writes the fix with that context, another agent runs some checks, and a PR opens with the risk already scored for the right reviewer. Most importantly, no person was responsible for pushing the work from step to step.
Now, that doesn’t mean that every step is deterministic. A factory combines two kinds of flows: deterministic and non-deterministic. Deterministic flows would handle the known paths like routing, gates and permissions. That’s because they are driven by rules the platform team has already defined, not the model's judgment that day. For example: automatic deployments are possible when they happen in a non-tier-1 service, with a low blast radius, while there are no open incidents, and when priority is not critical. All must pass, or else one failure goes to a human. A release on a Saturday is held by the deploy-window check and goes out only with an approved hotfix.
Other parts can be non-deterministic. Picture a builder agent that discovers halfway through a session that the feature has to support currencies with three decimals. Here, an agent pauses the data backfill, drafts the change to the spec, and asks the PM to approve it. The work goes back to the spec stage and flows forward again. Or imagine a release breaks checkout: the SLO check rolls it back (deterministic), then an agent diagnoses a cache stampede and puts the PR containing the fix at the top of the review queue. In both cases the agent had to choose the path on its own but also, brought in a person that can steer the session or take over the work.
The open question is how much the agent should decide alone. Too little and humans drown in escalations. Too much and review becomes a rubber stamp. Whatever the answer, the job changes. Once we get the balance right, approving agent work will only actually be a minor part of an engineer's week. Most of it will go into building and improving the factory itself: the flows, the standards, the context, and the rules for when an agent calls a human.
One more piece of the puzzle is where the agents run. Today, most still live on a developer's laptop, with their own context and their own rules. Software factories that work for large teams can’t be built from hundreds of developer’s local environments. But when agents run centrally, every agent a team creates can work from the same context and the same rules. Software factories must run in the cloud, and local is the exception.
Managing chaos requires two harnesses
Agents come with their own harness that turns the model into something useful. It makes sure the agent works towards the goal you’ve set for it. But unless you’re working solo, there’s another harness to be built: the platform or organization harness. It needs to answer questions like what should the agent see, what actions can it do, when can it act, and how should a human supervise it. The same agent would behave differently in two companies because that harness is different.
So the real job for platform engineering is not picking the best agent. It's building the layer that makes any agent behave well in your org, according to your unique standards.
Where software factories can start to break
First attempts at a full flow almost always work. That’s usually because someone gave a ticket to a coding agent and a PR came out the other end. Maybe it even handed it over to a code review agent too. But the trouble starts at hundreds of flows and thousands of agents, or even earlier. You start running into issues of scale. Which data can each agent see? Who owns each one? How much can each team spend? Who reviews hundreds of PRs a day? As you can see, running it at scale is hard.
Review is a good example. Say a feature lands as five PRs across five repos. If it gets reviewed one by one, nobody sees the whole change. But if it gets reviewed as one change, with a plain-language summary, a blast radius pulled from your service catalog, and a risk score, that’s much more valuable.
Agent skills are another example of something that’s easy to start and hard to scale. Skills are the reusable instructions agents load to do their work, and today, teams can add them without anybody reviewing or tracking them. A properly scaled system should block an unapproved skill and flag it right away, instead of finding out after a bad release.
Three ways to build an AI software factory
The first option is to assemble it yourself from separate tools. You retain full control, but your team needs to build and maintain a custom context layer, custom permissions, custom audit and metrics, and keep them all current as things change (and they will, a lot).
Second, you can buy a vertical product that gives you most of the factory in one box. It's fast to start, but it's built around the vendor's own agent, and nobody ships a complete version yet. The risk with this option is vendor lock-in as your model, agent, are harness can’t be easily swapped out in the future.
Or, you can build on a horizontal platform that connects to the tools you already run, and add any agents you want into the system. You still design your own flows, because your company's SDLC is unique.
Most teams think they need to make one decision: which agent to use. In reality, they have two decisions: which agent to use, and what foundation it should run on. Agents change every few weeks. The foundation should last years. A team can build or choose any coding agent, but it should be shaped around its own systems, and put to work on top of a shared platform with context, ownership, and policies. When a better agent or model ships, it gets swapped in without having to rebuild everything the agent depends on.
Not choosing is an option too. In that case, you’ll likely end up with each team building their own little, inefficient factory that works differently from the others and doesn’t follow organization standards.
Where to start
My advice is don't start with a grand design. Find out where agents are waiting on a person, on missing context, or on manual work. Suppose verification is where work waits longest and code owners are the least happy people on the line. That would mean that review is the bottleneck, not coding. A faster agent won't fix that. What would help are local fixes like ranking the review queue and approving related PRs as one change.
Every time you fix a manual part of your SDLC with an agent, make sure they use and strengthen the shared pieces: context, orchestration, controls. Do that long enough and they converge into a software factory that works for the whole team.
Yonatan Boguslavsky is co-founder and CTO of Port.