When the approved path is harder, developers build another one
The fastest route to shadow AI can begin with an AI policy itself. Developers face delivery pressure and ambiguous engineering problems, so a useful assistant has immediate value; when the sanctioned route is slow, vague, or detached from normal work, an engineer can paste sensitive material into a public model or quietly install an unapproved coding assistant. Stack Overflow, which sells products and services to developers and has a commercial interest in how they adopt AI, reports that 84% of respondents use or plan to use AI tools. At that level of demand, responsible adoption becomes a developer-workflow design problem.
That workflow problem also appears in Microsoft’s Work Trend Index. Microsoft, a major provider of enterprise and developer AI products, benefits commercially from broader workplace AI adoption, and its index found widespread bring-your-own, or shadow, AI use while many people were reluctant to disclose using AI for important work. Approval delays can make such behavior more likely when they clash with pressure to ship. A single process for uses with materially different data access, autonomy, and production reach can create similar friction because developers must resolve the practical conflict themselves.
That conflict shaped Ryan Donovan’s recent Stack Overflow podcast conversation with Microsoft’s Sarah Bird, where responsibility was framed around impact, accountability, and deliberate design of the workflow between people and AI. Stack Overflow’s broader findings add another constraint: developers are adopting AI while remaining wary of what it produces. More developers distrust AI accuracy than trust it, and their leading frustration is output that looks nearly correct but requires extra debugging. High demand and low trust together make workflow design central to responsible use.
The governance question consequently starts before anyone violates a rule. Leaders need to ask whether compliant behavior is fast, clear, and observable, whether controls reflect the risk of the task, and whether developers can challenge a result or escalate a problem within their normal workflow. A policy can define the boundary, but the engineering system has to make that boundary practical under delivery pressure.
Treat shadow AI as governance telemetry
Once unapproved use appears, it can show leaders where the approved workflow has failed. The use can still create serious security, privacy, or reliability risk, but an incident also provides evidence of unmet demand. Shadow AI can act as governance telemetry, meaning operational signals that show where the sanctioned process creates enough friction for people to work around it. That signal can turn a violation into input for improving the system while the unsafe use is still stopped.
Using that signal starts with the work developers were trying to complete. Teams can identify which tasks drive people toward outside tools, find the friction in the approved alternative, and determine whether the missing element is data, relevant context, an integration, or permission. That sequence distinguishes a necessary restriction from an avoidable process obstacle. It also gives platform and security teams concrete engineering work.
Those obstacles can guide the design of approved infrastructure. Stack Overflow has discussed monitored gateways and approved AI platforms as ways to let developers experiment through visible infrastructure, an approach aligned with its commercial interest in developer workflows and AI adoption. Such infrastructure gives an organization a place to apply controls while developers still get useful capabilities. An effective approved service gives them relevant context, straightforward access, understandable limits, and a rapid escalation route when ordinary rules do not fit the task.
Those access choices matter because excessive approval burden can increase delay without a corresponding safety improvement. Developers under pressure may then avoid the official process, conceal experiments, or choose tools readily available outside organizational controls. Governance consequently loses visibility where it most needs it. An investigation of shadow use should ask which part of the workflow made hiding or routing around the system seem worthwhile.
That diagnostic approach still lets an organization stop unsafe behavior immediately. At the same time, the organization can ask what the incident reveals about the sanctioned path and use the answer to adjust access, permissions, integrations, and escalation. Governance can then learn from observed demand before the same pressure produces another hidden workaround.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.
Put governance inside the engineering path
Fixing those workflow failures requires policy to become an engineering interface: a set of decisions developers can apply while doing the work. NIST’s AI risk-management framework organizes the task around four functions, Govern, Map, Measure, and Manage, treating governance as an ongoing operating activity. For an engineering team, that means classifying the use case, specifying permitted models and data sources, establishing tests, identifying the approval owner, and deciding what evidence must remain in the development record.
Those classifications have to answer questions when a developer encounters them. Policy should establish what data can enter an approved tool, which repositories and systems that tool may access, what review generated code receives, and which decisions remain with a human. It also needs a defined response for harmful, insecure, or unreliable output and a threshold at which an experiment becomes a production system. Routine work should give developers enough information to make these decisions without arranging a committee meeting.
Clear answers also reduce uncertainty during adoption. Stack Overflow’s technology-adoption findings identify security and privacy as developers’ leading reasons for rejecting a technology, a finding connected to the company’s commercial interest in developer technology use. Operational rules address that concern when they tell an engineer what can be done safely and what requires escalation. Governance then becomes part of the developer’s decision process under time pressure.
Those decisions need enforcement points where engineering actions become durable changes: repositories, pull requests, build pipelines, access systems, and deployment workflows. NIST’s secure AI development guidance extends secure software practices across the development lifecycle. Applied to enterprise AI, that approach can put approved model configurations under version control, restrict access according to role, scan prompts and outputs for secrets, retain logs for higher-risk uses, and require tests before AI-generated changes can merge.
Those enforcement points also make human review more specific. GitHub, which sells AI developer tooling and therefore benefits when developers can use AI coding systems successfully, recommends functional checks, verification against relevant context, dependency review, collaborative review, and automation where appropriate as part of human oversight for AI-generated code. These checks turn the loose instruction to “review AI output” into defined engineering work. A pull request can carry both the generated change and the evidence needed to decide whether the change belongs in the system.
Specific review matters because AI adds failure modes that ordinary code review needs to recognize. OWASP’s generative-AI risk categories include prompt injection, sensitive-information disclosure, supply-chain weaknesses, improper output handling, and excessive agency, where a system has more ability to act than its task safely requires. Those risks affect controls around inputs, outputs, dependencies, permissions, and actions. Compilation alone cannot establish whether generated code is safe to use.
The same lifecycle logic runs through different kinds of guidance. CISA’s secure-by-design principles place responsibility in product design, while NIST provides risk and development frameworks, OWASP identifies generative-AI failure categories, and GitHub turns human oversight into concrete review practices. Together, these materials support an engineering approach in which developers encounter governance through the systems where they request access, write code, review changes, run tests, and deploy software.
Make the fast path stricter as the risk rises
Once governance is embedded in those systems, easier approved access can coexist with strong control. A well-designed path can remove needless waiting while enforcing role-based access, secret scanning, logging, tests, human review, escalation, and production monitoring more consistently than an informal approval process. The key design question is how much control each use requires.
That amount should follow the use case. Asking an AI system to explain code has a very different risk profile from giving an agent write access to production systems, because the second case has greater autonomy and can directly change an operational environment. Applying the same approval burden to both consumes review capacity while giving insufficient weight to their different consequences. Risk scaling keeps the low-consequence case fast and reserves stronger scrutiny for actions capable of creating larger failures.
Four dimensions provide a practical basis for that scaling: data sensitivity, autonomy, reach, and reversibility. More sensitive data warrants tighter access and handling controls; greater autonomy calls for stronger constraints on what the system can do; wider reach increases the number of users or systems exposed to an error; and poor reversibility raises the cost of recovery. These dimensions give engineering and security teams a practical classification for each AI use.
That classification determines how controls tighten. Role-based permissions limit who can invoke higher-risk functions, secret scanning reduces accidental exposure, and higher-risk logging preserves evidence for investigation and review. Pre-merge tests and explicit human ownership constrain generated software changes, while escalation routes handle uncertain cases. Once AI behavior reaches production, monitoring adds another control, and review becomes progressively stricter as sensitivity, autonomy, reach, or difficulty of reversal increases.
Risk scaling also clarifies what human oversight means for developer tools. GitHub advises users to treat GitHub Copilot “as a tool rather than a replacement” and to review and validate its generated responses; GitHub has a commercial stake in Copilot’s successful adoption. The guidance keeps the developer accountable for the resulting engineering decision even when generation is fast. The workflow supports that accountability by supplying the context, tests, review steps, and permissions needed for practical validation.
Because the controls follow a known classification, a stricter high-risk path can still be efficient. Teams can design tests, preserve required evidence, assign the owner, and involve security or privacy specialists during development, before work reaches final approval. Speed then comes from removing uncertainty and repeated handoffs while maintaining the standard applied to consequential systems.
Give every AI workflow a human owner, and make failures reportable
Risk-scaled controls still need someone who owns the outcome. Terms such as assistant, collaborator, and agent can describe how software behaves, but organizational accountability remains human, so every AI use case needs a named owner who understands the intended result and has the authority to stop or change the process. That owner defines acceptable performance, decides where human judgment must enter, and ensures that a failure gets a response.
That ownership must persist because NIST’s Generative AI Profile places risk management across design, development, use, and evaluation. Product leaders own the business decision, while engineering leaders own implementation quality and security and privacy specialists own the controls relevant to their domains. Developers remain responsible for code they submit, reviewers for approval decisions, and operators for production monitoring and incident response. The division gives each stage an accountable person when consequences emerge.
Ownership works only when important signals can reach the owner. A developer may first see a problem as leaked context, a fabricated dependency, insecure-code encouragement, an odd completion, a suspicious package, an undocumented data path, or a plausible-looking result that contradicts domain knowledge. These signals can look minor before their cause or scope is known. Early reporting therefore becomes part of operational risk management.
That reporting depends partly on team conditions. Google’s Project Aristotle identified psychological safety as the most important dynamic in its study of effective teams, and Google characterizes such an environment as one in which people can take interpersonal risks, ask questions, and surface mistakes. For AI systems, that property has a direct technical consequence because weak signals become visible sooner when developers can challenge a tool, an assumption, or an approved process without expecting reporting itself to count against them.
DORA research similarly links psychologically safe software-delivery cultures with stronger performance and resilience. Managers can weaken that mechanism when adoption targets make questioning an AI tool look like resistance. After a failure, useful questions concern what failed, how the workflow encouraged the mistake, and which safeguard would reduce the chance or consequence of recurrence. Asking why a developer “trusted the AI” personalizes the failure and can signal that raising the next warning carries personal risk.
Because automated controls catch known patterns, engineers and domain experts still need a path for suspicious behavior that has no existing rule. A reporting culture supplies that path, turning observations into engineering action. Owners can then adjust permissions, tests, models, review requirements, or the workflow itself.
Train for the decisions developers actually make
Once ownership and escalation exist, training needs to prepare developers to use them during real engineering decisions. Stack Overflow’s developer trust analysis treats effective AI use as a skill involving prompt structure, evaluation of outputs, and integration of generated code into existing systems, within the company’s broader commercial interest in how developers use technology. That skill matters because developers already report both low trust in AI accuracy and debugging work from nearly-right output. Training therefore has to prepare people to verify results as part of the task.
That preparation should work more like a license to use the system than a generic awareness presentation. Developers need to know which tools are approved, what data those tools may receive, which common failures to expect, what review is required, and where to escalate an uncertain case. Organization-specific examples can then turn those rules into decisions resembling the work people will actually perform.
Those decisions differ by role, so the practice should differ as well. Backend engineers need experience reviewing generated database migrations, authentication logic, and dependencies; data engineers need to protect sensitive records and validate transformations. Engineering managers need criteria for determining when a pilot requires security, legal, or architecture review, while platform teams need to observe agent behavior and constrain permissions. Role-specific training places each decision with the group able to affect its risk.
Practice can also follow the way developers already learn. Stack Overflow’s learning findings show that developers build AI skills through multiple resources, including AI-enabled tools themselves. Organizations can use that behavior through sandbox exercises, deliberately flawed outputs, examples of acceptable prompts, and reusable internal patterns. Practice with errors is especially valuable because it develops recognition of the almost-correct results that create extra debugging and verification work.
Because practice ends, its useful guidance should remain in the engineering environment. Repository instructions can establish local rules, review checklists can make evaluation repeatable, reusable test suites can catch recurring failures, and approved prompt patterns and documented examples can reduce improvisation. Attendance records show that a session occurred, while these artifacts continue to shape later engineering decisions.
Measure team outcomes
Once those trained workflows are operating, adoption counts still cannot show whether they work well. Licenses assigned, prompts submitted, and weekly active users measure activity, while engineering value depends on quality, speed, risk, and the work carried by the whole team. A growing usage number can coexist with more correction, review, or production trouble, so measurement has to follow the engineering outcome.
The 2024 DORA research shows why broader measurement is necessary. Higher AI adoption correlated with improvements in documentation quality, code quality, and review speed, while the research also identified possible negative effects on software-delivery performance. Those mixed results mean individual productivity or adoption cannot stand in for system performance. The relevant unit is a defined workflow and its outcomes before and after AI enters it.
A defined workflow lets teams measure cycle time, escaped defects, rollback rate, security findings, review burden, documentation quality, incident volume, developer satisfaction, and time spent correcting AI output. The useful subset depends on the task and its risks, but measurement should follow the work far enough to expose transferred costs. Nearly-right output can save generation time while adding debugging, and a fast local change can create review or maintenance work elsewhere.
That downstream work is also visible in Stack Overflow’s survey of agents, from a company with a commercial interest in developer AI use. Agents indicate gains in individual productivity without corresponding gains in team collaboration. An engineer can therefore finish a task faster while colleagues absorb verification, integration, or maintenance effort. Responsible measurement includes that downstream work so an apparent local gain is judged against its effect on the engineering organization.
Following the work across those boundaries also makes the approved path itself measurable. Leaders can see whether useful context and straightforward access reduce the incentive for shadow tools, whether controls scale with data sensitivity, autonomy, reach, and reversibility, and whether reporting routes surface problems early enough for owners to act. Those outcomes show whether governance is functioning inside access, repositories, reviews, tests, deployment, monitoring, ownership, and escalation, the places where developers actually make consequential decisions.
The bottom line
Responsible AI adoption is ultimately an operating-model decision, not just a policy decision. Developers will use AI where it helps them deliver, and governance will be most effective when the approved path fits that reality. Leaders therefore need to make safe use easier to identify, access, and follow while reserving heavier controls for systems with greater sensitivity, autonomy, reach, or difficulty of reversal.
That requires coordination across engineering, security, privacy, product, and platform leadership. Shadow AI should inform where approved workflows need improvement; higher-risk uses should have clear human owners; and reporting, training, review, and monitoring should exist inside the systems where work already happens. Controls that depend on developers leaving their workflow to interpret policy are less useful under delivery pressure.
The executive measure of progress should not be AI adoption alone. Leaders need to know whether AI improves delivery without shifting hidden costs into debugging, review, security incidents, or maintenance. When organizations measure those outcomes and use them to refine the approved path, governance becomes part of how engineering operates rather than a separate process developers must work around.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.


