More AI-generated code does not guarantee more engineering velocity

Engineering teams report sharply different experiences with agentic coding tools, which can carry out multi-step coding work with limited human direction. Some engineers claim “x10 or x100 velocity” from tools such as Claude Code and Cursor, while others report code-review bottlenecks and senior engineers becoming “full time garbage collectors,” as productivity and enjoyment fall. Those experiences point to the engineering question that matters: what happens to the rest of software delivery when code generation gets faster?

The question matters because code generation is only one part of software delivery. A developer can produce implementation work much faster while the team takes longer to inspect it, correct it, test it, release it safely, or understand it after a failure. Higher coding throughput can push work downstream, so equivalent gains in engineering velocity depend on the condition of the software development lifecycle (SDLC), the process that takes work from requirements through development, review, release, and operation.

The SDLC makes AI readiness a property of the engineering system around the tool. A lifecycle with clear requirements, architectural constraints, effective review, meaningful testing, fast recovery, and outcome measurement can control an increased rate of change. When those practices are weak, AI can increase the volume of problems they already allow. The sharply different experiences can coexist because teams are adding generation capacity to different engineering systems.

AI multiplies the process it enters

Those different systems explain the multiplier effect. Strong practices can turn increased output into useful velocity because the organization can constrain what is produced and decide whether it is fit to ship. Weak practices can amplify existing shortcomings because more output reaches the same inadequate review, testing, and release controls. Engineering maturity therefore changes what acceleration means in practice.

The multiplier moves the important boundary downstream of generation. If an agent produces more code, somebody or some process still has to establish that the code solves the requested problem, follows the system’s conventions, behaves correctly, and can be operated safely in production. More output puts greater demands on those controls because more change has to pass through them. The useful question is whether the entire SDLC can absorb that higher rate of change.

The system-level question also changes how teams should evaluate AI tools. Coding speed by itself cannot establish whether delivery has improved when review queues grow, low-quality tests pass, or engineers struggle to diagnose production failures. Evaluation has to include both the constraints before implementation and the controls after it because generation sits between them. Useful acceleration depends on whether those surrounding practices can constrain, review, measure, and recover from the extra output.

That dependency explains why tool selection alone cannot settle the disagreement over AI productivity. Two teams can introduce agentic tools into very different environments and see different downstream effects even when both generate code more quickly. One team may detect and contain unsuitable output early, while another discovers problems after they accumulate. The readiness test consequently begins with the information available before an agent writes code.

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.

The readiness test starts before code generation

That information chain begins with requirements. If tickets routinely reach implementation with missing or incomplete acceptance criteria, an agent has incomplete information about what successful work means. It can head in the wrong direction, expand the scope, and produce something that does not meet user needs because the requested outcome remains unresolved. Faster implementation cannot correct a product decision that has yet to be made.

Because written tickets do not always contain everything needed for implementation, product clarification is the next constraint. A useful diagnostic is whether engineers frequently stop implementation to ask the product side for clarification. The interruption can be productive because the engineer has recognized an ambiguity before committing more work to it. If an agent proceeds without that clarification, an existing communication weakness can turn directly into incorrect or unnecessary implementation.

Once the outcome is clear, architecture constrains how the implementation should fit the system. Teams should ask whether architectural decisions are documented somewhere both engineers and AI agents can read and whether development conventions are written down. Without accessible conventions, an agent can make its own choices while generating an implementation. Those choices can produce functioning code while increasing inconsistency with the rest of the system, creating more work during review and later maintenance.

Together, these upstream diagnostics make AI readiness partly an information-quality problem. Acceptance criteria establish the outcome, product clarification resolves ambiguity, and architecture documentation constrains the implementation. An organization that holds critical knowledge informally asks an agent to operate without information that human engineers may recover through questions or experience. Increased output speed then magnifies the cost of unresolved requirements and undocumented design decisions.

Because those failures originate upstream, they can surface later as apparently poor generated code. An agent cannot reliably infer a product decision the organization has yet to make or a convention it has never recorded. Teams evaluating AI readiness need to inspect the inputs and constraints their current process provides before treating generated code as the primary unit of analysis. Once those inputs are adequate and generation accelerates, review becomes the next constraint.

Review and testing determine whether extra output is trustworthy

Review capacity becomes visible in pull-request size. Teams should ask whether their current PRs are small enough for a proper review in one sitting. Agentic tools can produce a large amount of code, increasing the risk that generated PRs become too large for careful inspection and are rubber-stamped instead. At that point, generation throughput has exceeded the team’s ability to make an informed judgment about what it merges.

When review falls behind generation, reports of bottlenecks and senior engineers cleaning up poor AI output become easier to explain. Large volumes of questionable changes can pull experienced engineers from their existing work into inspection and correction, while scope creep and implementation that misses user needs add more work. Generated changes produce value after the organization establishes that they belong in the product. Review capacity and quality therefore set practical limits on how much generated output the delivery system can use.

Even careful review needs testing because human inspection cannot establish every runtime behavior. A team that does not regularly run unit and regression tests, or does not treat quality as a shared responsibility, has fewer ways to detect problematic AI-generated changes. Increasing code volume also increases the amount of behavior that must be evaluated. The relevant readiness question becomes the strength of the testing culture already in place.

AI-generated tests make this dependency especially clear. Writing tests can be painful enough that engineers sometimes skip the work, so generating tests with AI can address a genuine weakness in the workflow. The intervention works when engineers can judge what was generated. A passing test confirms that its own assertions passed. Engineers still have to decide whether those assertions check the right behavior.

That judgment matters because poor generated tests can make a weak testing process worse: passing output can create confidence without establishing that the intended behavior has been checked. Engineers still need the technical competence to assess test coverage, assertions, and intended behavior even when AI does more of the writing. Automation changes who produces the artifact. Engineering judgment still determines whether to accept it.

The testing example reveals a broader constraint on AI adoption. A team may choose AI precisely because part of its SDLC is weak, but the surrounding organization still needs enough capability to evaluate whether the intervention improved that area. For generated tests, that means recognizing bad tests; for generated implementation, it means reviewing scope, behavior, and design. Without that evaluative capacity, more output can conceal the original deficiency.

Because pre-merge controls can still miss problems, trust also depends on what happens after release. Some failures become visible only in production, where a higher rate of change creates a different requirement. The organization must be able to recover quickly and measure whether delivery performance is improving.

Faster change only helps when failure is recoverable and measurable

Production failures make recovery speed part of the productivity calculation. Teams considering higher AI-driven throughput should ask whether they can roll back a bad release “in minutes.” AI increases the rate at which changes can be made, so slow recovery can turn that rate into a liability. A team that can reverse a failed deployment quickly is better positioned to contain the consequences of increased change.

Fast rollback contains a failure, while diagnosis depends on whether engineers understand what they shipped. A useful diagnostic is simple: can engineers explain the code they merge? If generated implementation enters production without that understanding, the same engineers may struggle to debug it when it fails. The productivity calculation has to include operational work created by code that was quick to produce but difficult to reason about under production conditions.

Once recovery costs are visible, measurement tells the organization whether faster generation adds up to improvement. Teams should establish whether they track change failure rate, meaning the share of changes that cause failures requiring remediation. Without that measure, an organization can increase AI use and code throughput while lacking a way to determine whether production outcomes have improved or deteriorated. Tool usage measures adoption; delivery outcomes establish its effect.

Those outcomes also set the boundary for return-on-investment decisions. A team needs measurable engineering results and reversible adoption to establish whether AI has helped in its own SDLC. Greater adoption supplies evidence about usage. It cannot by itself establish success, because success depends on what happens to the engineering outcome the intervention was meant to improve.

With an outcome defined, the decision becomes an engineering evaluation of consequences. If output rises while change failures rise and recovery remains slow, the additional rate of change is difficult to treat as useful acceleration. If a bounded intervention improves the chosen outcome and remains manageable when something fails, the organization has a stronger basis for continuing. Weak SDLC practices therefore give teams a reason to narrow the experiment.

Weak SDLCs call for narrower AI adoption

A narrower experiment lets a team use AI before every part of its SDLC is mature. The practical approach is to identify one aspect of the lifecycle that needs improvement and consider AI as one candidate for addressing it. The team can constrain that intervention instead of changing requirements, implementation, testing, review, and other stages simultaneously. A smaller scope limits adoption risk and preserves the ability to reverse course when the change causes problems.

The earlier testing example shows how that qualification works: AI-generated tests can address skipped test writing when engineers have enough testing competence to evaluate the result. The weak practice creates a reason to experiment, while the capability around it determines whether the experiment can be controlled. The same relationship applies elsewhere in the lifecycle, where the surrounding process has to make the intervention observable and reversible.

That adoption logic matters especially for teams with incomplete acceptance criteria, frequent unresolved product questions, undocumented architecture or conventions, oversized PRs, weak testing, slow rollback, absent change-failure tracking, or poor understanding of merged code. Broad AI adoption can amplify several of those weaknesses at once, so improving one SDLC aspect at a time makes it easier to identify what changed and contain adverse effects. AI may be the intervention for that aspect, but the decision should follow the problem being addressed.

These diagnostics are a set of dependencies rather than a complete maturity model. They connect increased generation to information before coding, judgment around generated work, and operational control after release, giving a team a way to bound an adoption decision around a weakness it can evaluate. Pluralsight extends this implementation and measurement approach in its guide, “Five steps for building AI-ready engineering teams and proving ROI.” Pluralsight has a commercial stake in AI-readiness guidance because it sells technology-skills and workforce-development products, so teams should weigh that guidance alongside the engineering outcomes they can observe in their own SDLC.

Main highlights

  • AI multiplies the process it enters: Faster code generation creates value when requirements, review, testing, and release controls can absorb the added change. Engineering leaders can assess AI productivity through delivery outcomes rather than coding throughput alone.
  • Readiness starts before code generation: Clear acceptance criteria, accessible architectural decisions, and documented conventions give AI agents the constraints needed to produce useful work. Product and engineering organizations can strengthen these inputs before expanding adoption.
  • Review and testing determine trust: Higher code output increases the burden on reviewers and testing systems, while generated tests still require engineering judgment. Engineering teams can keep PRs reviewable and verify that tests cover intended behavior before increasing AI-generated output.
  • Recovery and measurement determine useful velocity: Faster change increases operational exposure when rollback is slow or engineers cannot explain generated code. Delivery organizations can track change failure rates, recovery speed, and other outcomes to establish whether AI improves performance.
  • Weak SDLCs call for narrower adoption: Organizations with immature practices can target AI at one measurable SDLC problem rather than expanding it across the lifecycle. Bounded, reversible experiments make results easier to evaluate and adverse effects easier to contain.

Alexander Procter

September 30, 2026

11 Min

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.