Mark Phelps

Building · 10 Oct 2026 · 6 min · 1,346 words

When AI writes all the code

I delegate implementation to agents while staying involved in design, evaluating larger features, and checking that the whole application works as I expect.

You may have seen the term “software factory” lately in discussions about AI coding agents. The goal is to turn a specification into working software by automating implementation and validation.

I like building software. These days, agents write more of the code. I thought I’d care more about giving that up, but I don’t particularly miss typing every line myself.

I still like working out what to build, thinking through the architecture, and deciding how the pieces should fit together. And when an agent builds a substantial feature, I want to use it myself and check that the application as a whole behaves the way I expect.

I don’t need to personally exercise every small change. The question for me is how much implementation I can delegate while keeping a useful understanding of what I’m building and staying involved in the parts I enjoy.

What the software factory automates

The amount of human involvement varies. Some workflows keep people reviewing every change. Others aim to remove human code review entirely.

Dan Shapiro’s framework describes levels of automation from 0 through 5:

An inverted pyramid showing six levels from manual coding at level 0 to the dark factory at level 5, with decreasing human involvement in implementation.An inverted pyramid showing six levels from manual coding at level 0 to the dark factory at level 5, with decreasing human involvement in implementation.

Shapiro’s levels, shown as decreasing human involvement in implementation. Band widths are illustrative.

LevelHuman involvement
0You write the code, perhaps using AI for search or occasional autocomplete.
1You delegate discrete tasks, such as writing a unit test or adding documentation.
2You pair with AI, working through implementation interactively.
3Agents write the code; you spend your time reviewing their changes.
4You write and refine specs, review plans, and let agents execute before checking the results.
5The process becomes a black box that turns specs into software.

Shapiro calls level 5 a “dark factory”: a black box that turns specs into software. That describes how implementation happens. It doesn’t establish that people have to stop choosing what to build or setting architectural constraints.

StrongDM describes a software factory where specs and scenarios drive agents that write code and validate it without human code review. Their account includes validation scenarios stored outside the codebase to make it harder for an agent to rewrite a test just to get a passing result. In their description of the approach, they say: “Internal structure is treated as opaque.” They evaluate correctness through externally observable behavior.

That’s a meaningful change in how the implementation is checked. Building the validation system is still engineering work, though. A team can define requirements, set constraints, and decide what counts as a useful result while automating implementation and validation. Much of my own process has the same ingredients.

The decisions I want to keep making

I like coming up with requirements. I like thinking about the architecture, choosing dependencies, and working out how the pieces should fit together. Those decisions are a large part of why I enjoy building software.

I want to be able to explain why I chose an approach and what I expect it to do. For software delivered to a customer, I’d also want to understand what it depends on, what happens when those dependencies fail, and what evidence supports trusting it with their data. A working demo or a green test suite wouldn’t answer all of those questions for me.

Reading every line of code isn’t the only way to develop that understanding. Planning, documentation, tests, and using the application all contribute. They also have limits. I still need a way to check whether the implementation matches the design I think I’m shipping.

How I build with agents

Vibe coding is fun. I do it all the time. I’ve got five personal, mostly vibe-coded projects running while I type this. I’ve built two macOS apps, with a third underway, without knowing Swift. I’m also working on an Obsidian plugin whose code I haven’t read at all.

That last example is a limit on what I can claim. I can check whether the plugin behaves the way I want, but that alone doesn’t tell me whether its implementation is secure or maintainable. I’m willing to experiment that way on a personal project. I wouldn’t treat the same checks as sufficient for every piece of software I might deliver to someone else.

Here’s the process I use:

  1. Decide what problem I’m solving. I brainstorm what I want the app or plugin to do and what would make it useful.
  2. Work through the architecture. I think about the language, stack, dependencies, and how the pieces should fit together. The platform narrows some of those choices, but there are still decisions to make.
  3. Refine the requirements. I write a product requirements document (PRD) and go several rounds with different models until the plan is clear enough to implement. This is where I spend time arguing with the proposed approach.
  4. Set standards for the agents. I put project rules, linting, anti-slop hooks, test harnesses, and continuous integration (CI) in place so the agents have constraints and feedback.
  5. Evaluate substantial features myself. I check that the tests pass, but I also want to run the application and try larger features myself. I check the application as a whole to see whether it works the way I expect. The agents wrote the tests too, so I don’t treat passing them as the whole verdict.
  6. Keep checking after implementation. I add error reporting, logging, performance tests, and other observability where appropriate.

Some of that resembles Shapiro’s level 4. Some could fit a level 5 factory. I don’t think choosing a level tells me much about whether I’m getting what I want from the process. For a small change, I may be happy to rely on automated checks. For a substantial feature, I want to spend time with the result.

Checking whether I asked for the right thing

Automated tests check the scenarios we’ve defined. Using a larger feature myself also helps me decide whether those scenarios captured what I actually wanted.

A feature could satisfy the spec and still feel awkward to use (I run into this all the time). Two features could work separately but fit together badly. Or I could discover that I asked for something that doesn’t solve the problem as well as I expected. That happens more often than I’d care to admit. In those cases, the requirements need another pass, even if the implementation did what they said.

A factory can have ways to evaluate those things too. I don’t think personal testing belongs exclusively to level 4, or that automation rules out product judgment. I want to participate in that judgment myself, especially when the work changes how the application behaves.

That doesn’t make my evaluation a guarantee. Running the app won’t uncover every bug, and it won’t establish that the implementation is secure. I need other checks for those questions. But I do want to know what it feels like to use the thing I’m building.

Where I want the process to take me

I expect to delegate more as the tools improve. I may become comfortable automating changes that I’d want to inspect closely today. For a personal experiment, I might accept more uncertainty and keep going. For something a customer relies on, I’d want stronger evidence and a clear way to diagnose and recover when it fails.

I don’t need to commit to staying at level 4 to make those decisions. I need a process that helps me understand what has been checked, what remains uncertain, and where my own attention would be useful.

The time I save on implementation gives me more room to think through the design and use the application. I can try different approaches, apply my own taste, and make software that feels personal. Those are parts of building software I want to keep doing. When a larger feature is ready, I want to try it, see how it fits with the rest of the product, and decide whether it does what I had in mind.

Maybe I’ll end up using something I’d call a software factory. I still want to know what I’m shipping and enjoy the work that goes into it.

The letter

The next essay by email. It costs nothing.