AI coding agents can sharply reduce the cost of creating code, yet that gain can move the engineering constraint to a later stage. Once agents produce code faster than engineers can understand and diagnose it, verification, debugging and incident investigation begin to determine throughput. For engineering leaders, code-generation speed is therefore an incomplete measure of productivity because generated code still has to survive the rest of the software-delivery process.

That downstream constraint gives a precise meaning to the provocative question “Are AI coders more trouble than they’re worth?” The evidence supports a narrower conclusion: faster generation can create a bottleneck when human analysis cannot keep pace. Whether coding agents improve total engineering productivity remains a broader question, so the useful unit of analysis is the full path from generated code to software that teams can understand, verify and operate.

AI coding agents can create a verification and debugging gap

Evidence for that bottleneck comes from a poll conducted by independent research firm Coleman Parkes. It covered 300 software developers and engineering leads in the UK and US and was published in Undo’s report, Overcoming the limitations of coding agents in complex software systems. Undo provides debugging tools, so it has a commercial interest in the importance of debugging and in approaches that make debugging more effective. Most software development teams polled said AI agents struggle to identify issues in complex code, placing the difficulty where generated output has to be understood and corrected.

That broad finding becomes more useful when separated into forms of downstream work. Some results concern incorrect diagnosis, others incorrect generated code, while another measures time lost analyzing AI output. Together, they show several ways rapid generation can create additional work later in software delivery.

Poll finding Reported share
Teams that experienced AI hallucinations leading to incorrect diagnoses of code problems 93%
Teams saying agents introduce incorrect code too frequently, creating rework and delivery delays 55%
Teams losing productivity from analyzing AI-generated code at least monthly 94%
AI-generated code reaching production before teams fully understand what it does 35%

The 93% result matters because hallucination has a specific consequence during debugging. An agent can provide a wrong explanation of an existing problem, potentially directing an investigation toward the wrong cause. Engineers then have to verify the diagnosis before trusting the proposed fix, which adds another decision to the debugging workflow.

That diagnostic burden can begin earlier when the generated implementation itself is wrong. Fifty-five percent of respondents said agents introduce incorrect code too frequently, with rework delaying delivery cycles. Code may still be produced cheaply and quickly, but engineers give up part of that gain when they must identify bad output, determine the intended behavior and redo the affected work.

The need to inspect generated output extends beyond cases requiring rework. Ninety-four percent acknowledged losing productivity at least once a month because they had to analyze AI-generated code. The result shows that understanding machine-produced output consumes engineering attention for nearly all respondents, while the poll’s 300-person UK and US sample defines the population behind that finding.

Undo connects that analysis burden to how engineers spend their time. The company says developers tend to spend twice as long debugging as writing code, with debugging averaging 16.9 hours per week. Undo attributes that burden to software engineers being unable to keep pace with the coding agents they use, an interpretation aligned with the commercial problem its debugging tools address.

That pressure becomes more consequential at the production boundary. The poll found that 35% of AI-generated code reaches production before development teams fully understand what the code does. Coding errors can consequently enter running systems while the team’s understanding remains incomplete, turning comprehension of the generated implementation into part of an incident investigation.

These measurements establish a throughput mismatch within AI-assisted development. Agents can make code creation cheaper while engineers incur costs in analysis, rework, debugging and production investigation. For engineering leaders, productivity measurement therefore has to continue beyond generation and account for the work required to understand and operate what agents produce.

The hardest failures are the ones that look almost right

That need for understanding grows when generated software behaves plausibly while remaining subtly wrong. Greg Law, founder and CEO of Undo, draws the distinction this way: “When code is obviously broken, the cause is usually easy to find,” followed by the harder case: “Where engineers struggle is with code that’s almost – but not quite – right. Those are the times they lose days trying to unravel what went wrong and why.”

Nearly correct code creates a diagnosis problem because engineers first have to establish which behavior is wrong and then explain its cause. An obvious failure can quickly narrow an investigation, while subtle failures may preserve enough expected behavior to leave several possible causes in play. Law’s claim is that these cases can consume days as engineers work out both what failed and why.

That investigative work remains necessary when an AI system lacks enough evidence about the running application. Undo warns that applications can be difficult to explain and that AI systems can produce confident answers despite incomplete evidence about application behavior. Under those conditions, an agent can hallucinate a cause for a bug or an instability, so engineers have another diagnostic output to check before acting on it.

The dependence on evidence separates code generation from debugging. Generating a plausible implementation begins with a requested change, while diagnosing unexpected behavior requires enough information to reconstruct what the software actually did. When an agent lacks that information, confidence in its answer cannot establish that the causal explanation is correct.

The resulting human workload shifts toward downstream problem-solving when an agent cannot establish the cause reliably. Undo says this is where engineers now spend most of their time. Producing a first implementation may require less effort, while determining whether generated software behaves correctly and investigating unexpected behavior requires more judgment.

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.

The full AI coding workflow shows where productivity gains can stall

The difference between generation and diagnosis becomes clearer when both are treated as stages of one workflow. An AI agent can rapidly generate code, after which engineers have to understand and analyze what it produced. Some output can proceed into production while that understanding remains incomplete, allowing coding errors or unexpected behavior to emerge in a running system.

A production failure then makes the earlier comprehension gap part of the investigation. An AI agent working from incomplete evidence can misdiagnose the cause, which means an engineer may have to verify its explanation, take over debugging and investigate the incident. Rework and investigation can consequently consume part of the productivity gained by accelerating the initial implementation.

That sequence can change where engineering capacity is spent even when the agent is highly effective at producing code. Human comprehension and debugging may grow more slowly than generated output, so increasing generation volume raises demand on the stages that follow it. Once that happens, practical throughput depends on the team’s capacity to establish what the code does, whether it is correct and why it behaves unexpectedly.

The authors of the Undo report argue that continual human intervention cannot sustain that pace. If a person has to guide AI through every bug fix or determine the cause of every unexpected behavior, human investigative capacity remains the limiting resource while agents continue generating code rapidly. Their argument places understanding, verification, debugging and incident response inside the productivity calculation that begins with AI-generated code.

The proposed next step is AI that can investigate

That capacity constraint leads to Undo’s proposed response: apply AI across more of the software-delivery cycle and give debugging agents the ability to investigate. For debugging, that means an agent that can independently gather evidence and data and use them to reach accurate conclusions. An agent investigating production behavior would then have information about what the software actually did while running, giving it evidence for a causal explanation.

Runtime context is central to that proposal because detailed evidence about executing software can provide a stronger basis for determining what happened and why. That capability directly targets the hallucinated-diagnosis problem identified in the poll. For engineering leaders, the resulting evaluation question is whether investigative capacity can grow with generation capacity once agents can produce code faster than teams can establish what it actually does.

Law frames Undo’s proposed response explicitly around this imbalance: “Their challenge is that while agents are great at writing reams of code quickly, they’re less capable at debugging it. The result is engineers are being buried in an avalanche of code that’s well beyond human capacity to debug. That’s why we have to give them a way to make AI better at debugging, by feeding agents with the rich context of what code actually does at runtime.”

Law’s prescription comes from a company that sells debugging tools, giving Undo a commercial benefit if engineering teams treat debugging and runtime evidence as priorities. The mechanism behind the proposal remains concrete: debugging requires evidence about application behavior, and Undo proposes giving AI systems access to that evidence so they can investigate more of the work themselves. An AI system that can independently acquire and reason over runtime evidence would address the investigative stage where Undo argues rapid code generation is creating a capacity constraint.

Key takeaways for decision-makers

  • Measure the downstream workload: AI coding agents can accelerate code creation while increasing analysis, rework and debugging. Engineering leaders can assess productivity across the full delivery cycle, including the cost of understanding and operating generated code.
  • Prepare for subtle failures: Nearly correct AI-generated code can be harder to diagnose than obvious failures, especially when agents lack runtime evidence. Engineering teams can require verification of AI diagnoses before fixes reach production.
  • Track the full AI coding workflow: Faster generation can make human comprehension and debugging the constraint on delivery throughput. CTOs can monitor verification time, rework and incident investigation alongside code-generation gains.
  • Give debugging agents runtime evidence: Undo argues that investigative AI needs access to evidence about how software actually behaved to diagnose failures reliably. Platform and engineering teams can evaluate whether runtime context helps agents investigate bugs with less human intervention.

Alexander Procter

October 5, 2026

8 Min

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.