AI coding can make software appear faster than engineering organizations can decide whether it is safe to ship. One Developer Survey found a clear divergence between AI use and confidence:
| Measure | Earlier figure | Later figure |
|---|---|---|
| AI usage | 76% | 84% |
| Trust | 40% | 29% |
Those Developer Survey figures show that adoption and confidence can move in opposite directions. Coding agents can solve problems faster and produce whole applications in a fraction of the previous time, but generation speed measures only one part of engineering productivity.
That gap matters because agentic coding tools, meaning AI agents that can perform multi-step coding work, reach across more of the software development lifecycle (SDLC) than conventional development tools usually did. As generation becomes cheaper, scarce engineering effort moves toward specifying what should be built, providing the right context, validating large outputs, coordinating with other people, reusing proven work, and accepting responsibility for production behavior. AI coding creates durable value when teams redesign these trust-producing processes around its speed.
AI adoption is rising even as trust falls
The divergence between use and trust changes how engineering leaders should interpret rapid AI adoption. A tool can spread because it gives individual developers powerful generative capacity while the complete path from requirement to reliable production software remains difficult to trust. Producing an application quickly proves that generation has accelerated. Review, deployment, operation, and maintenance still have their own constraints.
Those lifecycle constraints matter more as agents take on work throughout the SDLC. A conventional tool generally has a bounded job that developers learn through repeated use, while an agent accepts natural-language instructions, generates substantial amounts of code, and can handle tasks across domains that previously involved several tools and people. Teams therefore have to measure how quickly generated work becomes software they are willing to own.
Developer trust comes from predictable workflows
Willingness to own software depends partly on predictable tools and processes, and engineering teams have traditionally built that predictability through repetition. Developers use a tool, become proficient with it, customize it around their work, learn its limits, and observe whether the same operation produces consistent results over time. The useful test is mundane: experience on Monday should remain relevant on Tuesday. Repetition turns explicit knowledge into muscle memory and makes tool behavior easier to anticipate before a developer acts.
An argument over Vim and Emacs about six years ago exposed how valuable that accumulated experience can be. A provocative piece about increasingly powerful IDEs questioned why anybody would still use Vim or Emacs and described their users as working like “cavemen.” The comments became hostile and defensive on both sides, including criticism that the argument misunderstood developers and advocacy from committed Emacs users. Beneath that conflict was an explanation for why apparently old tools could remain extremely productive.
For a novice, Vim and Emacs can be unintuitive terminal applications whose commands and keystrokes require substantial memorization. Expert users can reach a point where operating them feels almost like thinking because years of repeated use have removed conscious effort from common actions. Extensive customization also lets an individual adapt the editor closely to a personal workflow. David Thomas and Andrew Hunt, co-authors of The Pragmatic Programmer, describe developers as needing “sharp tools”; expertise makes a tool increasingly fit the developer’s established way of working.
That accumulated proficiency explains why moving to a more capable tool can initially slow an experienced developer. Tricia Gee, a developer productivity advocate, said, “One of my reticences for embracing some of the AI programming is because I’m faster with my IDE because I know how that works.” Moving from a terminal workflow to an IDE, or from an IDE to an agent, means changing a process that already produces useful results. Feature training alone cannot preserve years of practiced behavior.
Gee has seen the same resistance among expert Vim and Emacs users presented with IntelliJ IDEA’s refactoring features. “I’ve seen the same thing when I’ve worked with people who know Vim and Emacs very well. You could use the refactoring tools in IntelliJ IDEA. They’re like, yes, but that requires a learning curve. I’ve spent so long using whatever tool it is. My fingers know what to do. Understanding a tool very well, no matter what it is, becomes a lot of unconscious competence.” Tool migration can destroy some accumulated speed before its new capabilities create value.
AI makes comparable predictability harder to build because agent capabilities are changing quickly and their outputs are less predictable. Developers can generally learn the bounded role and limits of an IDE, containerization tool, or static analyzer. Linters, automated unit tests, and continuous integration and continuous delivery (CI/CD) systems likewise perform understood functions inside a larger workflow. Probabilistic agents enter more parts of that workflow while changing what they can do, so developers have less stable experience from which to infer future behavior.
Even mature tools derive their engineering value from the process around them. Better IDEs can improve implementation work, CI/CD systems can improve delivery, and issue trackers with story points can support estimation, while culture, habits, decisions, and coordination determine how well those capabilities translate into outcomes. Sizeable engineering organizations can reject capable tools when they conflict with how people actually work. Technical capability creates value after it fits, or successfully changes, the surrounding process.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.
When generation gets cheap, review and validation become expensive
That surrounding process comes under immediate pressure when agent output grows faster than the capacity to validate it. An old joke held that changing 100 lines was the way to get a pull request approved quickly. A coding agent can now change hundreds of lines almost instantly and place the resulting diff in front of a human reviewer. Generation has become much cheaper while the reviewer still has finite time and attention.
Finite review capacity can make code review the new bottleneck even when writing code becomes dramatically faster. Large diffs require reviewers to understand behavior, interactions with the existing codebase, edge cases, and potential production effects. When generated volume exceeds review capacity, teams face pressure to accept increasingly large changes with increasingly shallow inspection. Rubberstamping restores throughput on paper while weakening the approval mechanism itself.
That review constraint has encouraged automated code review and LLM-as-a-judge systems, in which a large language model evaluates another system’s output. These systems can scale parts of validation, but relying on an AI review of AI-written code creates its own trust decision. The review layer needs controls, evaluation, and accountability appropriate to the decisions it makes. Automation can relocate validation work while trustworthy validation remains necessary.
The growing tooling around agentic development reflects how many lifecycle constraints have moved at once. AI SREs are being introduced into operations, automated reviewers into pull requests, memory and context managers into agent interactions, and improved control planes and harnesses into systems that constrain agent behavior. These additions can strengthen an AI-enabled engineering organization. Their proliferation also shows that faster generation changes several parts of the SDLC together.
Those lifecycle effects become concrete after software reaches production. Cloud-native applications consume compute, memory, and network traffic, while hosted dependencies and APIs create continuing costs. Failures add harder-to-budget costs through downtime, security breaches, and missed opportunities. An agent can make code cheap to create while those operational consequences remain, so engineering productivity has to include the lifecycle that follows generation.
Because production consequences persist, existing engineering controls still matter even as agent workflows change them. Linters, tests, and CI/CD encoded parts of earlier engineering processes, and agentic workflows may need to adapt those controls to the amount, origin, and scope of generated work. Improving individual tools helps when organizational practices evolve with them. The same issue becomes even more important when agents begin crossing organizational as well as technical boundaries.
Agents can bypass the collaboration that used to contain risk
Those organizational boundaries traditionally distributed knowledge and authority among different roles. Product managers developed feature and function requirements from research, customer conversations, and knowledge of competitors. Architects applied experience and understanding of the existing technology stack, engineers implemented and reviewed the code, and QA tried to break what had been built. After deployment, DevOps and SRE teams monitored production performance and resource consumption.
That sequence made different kinds of expertise visible at different stages even when the process itself was imperfect. Access controls, audit trails, diffs, and CI/CD checks also constrained what a single person or tool could change without scrutiny. These mechanisms reduced the blast radius, meaning the scope of damage one mistake could cause, by distributing authority and recording consequential actions. Trust came partly from knowing who had made a decision and which controls it had passed.
That ownership remains with people when AI enters the workflow. Charity Majors, CTO of observability company Honeycomb, has a commercial interest in engineering practices that improve the operation and diagnosis of production systems, and she rejects language that makes human involvement sound secondary: “Human-in-the-loop sounds like a pity invite. I made the loop, I own the loop, I’m the only reason that loop exists. It is MY f***ing loop!” In practice, the developer who pushes a commit remains responsible for it, and the people who approve a pull request remain responsible for that approval. When a Friday deployment broke production before AI, responsibility stayed with the people and process around the deployment; an agent cannot assume that accountability.
Accountability becomes harder when an agent lets one developer cross boundaries that previously triggered consultation. Jaime DeLanghe, chief product officer at workplace collaboration company Slack, benefits commercially when teams see shared communication as important to software work. She described the risk: “You don’t have to, say, check in with your designer about the design because you just got an agent to build the design, and you don’t have to work with another engineer in a specific domain to understand their code base because you just ask the agent the question. You could go way down this path and create a massive PR.” The individual gains speed while the organization can lose interactions that surfaced assumptions and specialist knowledge.
Losing those interactions creates the possibility of a “silo of one.” A developer working through an agent can move across requirements, design, implementation, unfamiliar code, and deployment with greater independence, which means fewer people necessarily see the reasoning before the work becomes a large diff. The resulting risk is lost shared context and less visible decision-making. A massive pull request at the end cannot recreate consultation that would have changed an earlier decision.
Because the missing information is collaborative, some teams are moving agent interactions into shared workflows. One proposal is to prompt the agent in a shared room where colleagues can comment on and modify the developing approach; another is to attach prompt transcripts to pull requests. Dane Knecht, CTO of internet infrastructure company Cloudflare, has a commercial stake in tools and infrastructure used to build, deploy, and operate software, and sees value in that visibility: “It’s so cool to be able to open up a PR and actually see how the developer was thinking and how they were problem solving.” A transcript can expose reasoning and instructions that the final diff alone cannot show.
Trust has to be engineered before generation
Shared transcripts improve visibility after work has begun, but the same reasoning points to an earlier control: the information supplied before code exists. Teams need to state intended behavior, provide verified organizational context, and reduce ambiguity before asking a model to generate. Human review can then evaluate an output produced under clearer conditions. Reviewers spend less effort reconstructing requirements that should have shaped the implementation from the start.
Natural language makes that preparation difficult because conversational ease differs from specification precision. Bjarne Stoustrup, creator of C++, put the distinction directly: “Code is a precise statement of a solution,” while “English is a lousy language for expressing things that have to be unambiguous.” A coding agent receives instructions in a medium where a requirement can be incomplete or open to interpretation. Fast generation raises the cost of unresolved ambiguity because a model can turn it into substantial implementation before anyone notices.
Specifications therefore become an active engineering control. Teams can use prompts or spec.md files to state required behavior before generation, including constraints that a developer might previously have handled implicitly while writing code. Scott Hanselman, VP of Developer Community at Microsoft, whose employer sells developer platforms and AI-assisted development tools and benefits commercially from effective AI development workflows, summarizes the risk as: “If you leave anything up to chance, it will be left up to chance.” The practical requirement is to expose assumptions that an agent cannot reliably infer.
Hanselman’s ring-light application shows how small those omitted assumptions can be. He was working on x64 but wanted the application to support ARM, so he had to tell the agent, “Make me an ARM version.” Without that instruction, the ARM version would not have happened. The engineering change is that a requirement held in a developer’s head has to become explicit input before the agent can reliably act on it.
Explicit requirements still cannot contain all the organizational knowledge that software work depends on. Senior developers accumulate conventions, previous decisions, local constraints, known failure modes, and relationships among systems over time. Repeating all that material for every task would be expensive, and people would still omit relevant details. The context system around an agent therefore becomes part of the engineering system.
Agent systems can address that problem through “long-term and short-term memories,” meaning stored knowledge retained across work and task-specific information available for the current interaction. Useful memory requires a controlled sequence: an organization captures knowledge held by senior developers and the company, verifies it, stores it where systems can retrieve it, and supplies the relevant pieces when a task needs them. Stack Internal is one example aimed at capturing, verifying, and serving this organizational knowledge. Its role is to supply appropriate trusted context for the task.
Once generation starts from explicit requirements and verified context, peer review remains part of the process. Passing review and reaching production then creates another opportunity to retain what the organization has learned. Proven functionality can become discoverable and reusable for later work. Regenerating an already successful component introduces variation and discards some of the confidence earned through review and production use.
Laly Bar-Ilan, Chief Scientist at component-platform company Bit, has a commercial interest in software reuse and component sharing, and frames this tension through established software practice: “The DRY (don’t repeat yourself) principle is our main principle,” followed by “AI today is inherently WET (write everything twice).” Here WET means “write everything twice,” a pattern in which cheap generation can make another implementation feel easier than finding and integrating an existing one. That local convenience can create duplication across the organization.
Bar-Ilan gives the concrete case of UI components. One developer asks an agent to generate a button, then a developer on another team makes the same request and gets another button in the codebase. Both interactions can look individually productive while creating duplication at the organizational level. An agent workflow that supports reuse needs access to existing components and the organizational knowledge that tells it when reuse is the established path.
Reuse makes the process cumulative because later requests can benefit from earlier review and production experience. Responsible humans retain accountability for decisions, AI contributions stay visible, and teams can share and revise working practices as they learn. Specifications and verified context shape generation; peers then review the result, production use supplies feedback, and successful components become available for later tasks. Each stage preserves information that would otherwise be lost before the next request.
Encoding those practices into tooling follows an approach engineering teams already use. Linters encode rules that developers would otherwise check manually, automated tests encode expected behavior, and CI/CD systems encode deployment checks and sequences. Agent workflows need mechanisms suited to a wider scope of action and probabilistic output. The organizational move remains familiar: identify a practice that creates confidence, make it explicit, and make the normal workflow execute it consistently.
The right AI workflow also knows when to avoid AI
Encoding an agent workflow still leaves a prior question: whether the task benefits from AI at all. Large language models produce probabilistic, or non-deterministic, outputs, meaning repeated requests can produce different results instead of the fixed behavior expected from deterministic code. That variability is useful for work that benefits from generation and flexibility. Tasks already solved reliably by deterministic implementations gain less from introducing it.
Anil Dash identifies that tool-selection error directly: “We are trying to apply non-deterministic systems to a lot of scenarios where you should have deterministic code.” His practical counterexample is equally direct: “The humble bash script that has been running for six years is fine.” Six years of reliable behavior can be more valuable for that task than generative flexibility.
That example makes tool selection a question about the work and its full cost. Teams can use probabilistic generation where it provides meaningful gains, such as producing substantial new implementation quickly, while retaining deterministic code for tasks it already handles reliably. The earlier Vim and Emacs experience points to the same decision at the individual level: an experienced developer may remain faster with a deeply learned IDE, Vim, or Emacs workflow when switching consumes more time than it returns. Agentic coding earns its place where its gains survive the costs of specification, validation, operation, and process change.
Main highlights
- Treat trust as an adoption constraint: AI coding use can rise while developer confidence falls. Engineering organizations need predictable workflows that make agent behavior, limits, and ownership clear.
- Protect validation capacity: Faster generation can turn code review into the bottleneck. Engineering managers can control generated change size, strengthen automated checks, and match output volume to the organization’s ability to validate it.
- Preserve collaboration around agent work: Coding agents let developers cross requirements, design, implementation, and operational boundaries with fewer consultations. Shared agent interactions, prompt transcripts, and existing approval controls keep specialist knowledge and decisions visible.
- Engineer trust before generation: Product and engineering teams can make requirements explicit and give agents verified organizational context before code is produced. Peer review, production feedback, and reuse of proven components then make each successful implementation more valuable over time.
- Choose AI according to the task: Probabilistic generation creates value where flexibility and rapid implementation outweigh the costs of specification and validation. Platform and engineering teams can retain deterministic code and established tools where they already deliver reliable, efficient results.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.


