Software quality with coding agents: generate, understand, verify.
Greater code production capacity needs a process that can verify it. Measure the benefit through useful, maintainable software that reaches production.
The evidence needs context
Productivity and quality do not have an automatic relationship.
Borg and colleagues' experiment, published in 2026 with 151 participants, finds no significant differences in subsequent maintainability of code developed with AI assistants. The tasks took place in late 2024, before today's coding agents.
Liu and colleagues instead analyse around 302,600 AI-attributed commits in 6,299 repositories. They identify issues through static analysis, of which 22.7% remain at the latest observed revision. This is evidence of persistent issues in the sample, not a measure of every defect or a comparison with every human developer.
GitClear observes more duplication and less refactoring during the spread of AI. These are signals to monitor; alone, they do not establish a single cause. Reading these studies together suggests measuring effects within your own process.
Verification can become the bottleneck
If an agent changes more code than a team can understand and verify, the initial speed can become review and maintenance work. This is the risk I consider central.
My proposal is to separate production from validation: acceptance criteria and reference results should also be checked outside the process that generated the change. A second agent can help, but may share the first agent's mistakes.
- Behaviour, integration, contract and regression tests proportionate to risk.
- Static analysis, security checks and dependency checks.
- Reviews with explicit criteria and comparison against domain requirements and constraints.
- Changes small enough to make mistakes and consequences visible.
Not every line needs the same human scrutiny. There must, however, be a person responsible for the decision, with stronger checks where an error has greater consequences.
The harness must make errors visible
A harness contains context, rules, tools and procedures. Improving it can make an agent more effective; it can also propagate a mistaken interpretation faster.
I therefore include action limits and stopping conditions in the harness: failed tests, ambiguous requirements, unexpected differences or missing checks. Detecting an incorrect change before it reaches the product matters.
Maintain an understanding of the system
I use “cognitive debt” to describe a risk: the software still works, but fewer people understand why it was built that way. This is a useful concept, not an established standard metric.
Documented decisions, domain discussions and checks with the team help retain that knowledge. A developer's role includes requirements, architecture, context and judgement about changes; delegating execution still requires these skills.
Shopify: architectural trade-offs change
In September 2026, Shopify describes moving from React Native to Swift and Kotlin. It does not call the framework a failure: coding agents changed the cost of maintaining two native implementations.
The Shop app was rebuilt and released in about 12 weeks. Migration of the Shopify app, with more than 300 screens, is described as ongoing. This is one company's experience, not an estimate transferable to other applications.
For the migration, Shopify uses Helix: small checkpoints with tests, visual comparison and two independent agent reviews. Engineer approval is the default, although an autonomous mode is available.
The case supports a possibility: reducing implementation costs can make a previously expensive architectural choice viable. The outcome also depends on the process, tests and change controls.
A system that is easier to verify
Shopify also describes agent-addressable architecture: logic separated from the interface and fast tests through the command line.
From this I draw a broader proposal: clear responsibilities, explicit contracts, fast tests and understandable observability make a system easier for both people and agents to work on. Applying these criteria does not require the same stack.
Coding agents on legacy applications
A coding agent can work more independently when it has a verification loop: it changes the code, runs checks, reads errors and corrects them. Automated tests are central to this loop, alongside compilation and other relevant checks. A strict TDD process is not necessarily required; what matters is repeatable feedback on the required behaviour.
In some legacy applications, checking a change requires starting the entire system, a database with suitable data, external services and configurations that are difficult to recreate. When the logic is tightly coupled to these dependencies and automated tests are missing, setting up a verification loop takes additional work. A recent application can have the same constraints.
A database does not prevent automation: test environments and resettable data can be prepared. It may nevertheless be necessary to document startup, introduce tests describing existing behaviour and gradually separate some dependencies. This work belongs in the initial adoption assessment. Where verification remains mainly manual, I would allow the agent less autonomy and keep changes small, with explicit human checks.
Measure work the team can maintain
AI-generated code percentage, tokens, prompts or pull request counts are insufficient measures of benefit. I compare completion and review times, defects, regressions, rework and the ability to evolve the product.
The SPACE framework treats productivity as multidimensional. In that spirit, I propose assessing outcomes, collaboration and sustainable working practices together, without turning a single measure into a ranking of individual developers.
The objective I would choose is correct, useful and maintainable software at a sustainable cost.
Supporting sources
The studies use different methods and cover different settings. The proposals on validation and the developer’s role are working principles drawn from this evidence, rather than promises of a universal outcome.
- Borg et al. — Echoes of AI (opens in a new tab): experiment on subsequent maintenance; data collected in late 2024.
- Liu et al. — Debt Behind the AI Boom, version 2 (opens in a new tab): AI-attributed commits and issues detected through static analysis.
- GitClear — Maintainability Gap (opens in a new tab): observational signals on duplication, reuse and maintenance.
- Shopify — choosing native development (opens in a new tab): reasons for the change and agent-addressable architecture.
- Shopify — Shop app migration (opens in a new tab): experience and timings of that particular project.
- Shopify — Helix (opens in a new tab): checkpoints and verification for the Shopify app migration.
- Forsgren et al. — SPACE (opens in a new tab): complementary dimensions of developer productivity.
- Claude Code — Best practices (opens in a new tab): criteria and checks the agent can execute.
- Martin Fowler — Legacy Seam (opens in a new tab): separating dependencies to test existing code.
- Microsoft — Choosing a testing strategy (opens in a new tab): real-database tests, isolation and the limits of test doubles in EF Core projects.