The security boundary changes when AI can act
A model that returns a bad answer creates one class of security problem. An AI agent that can turn a bad decision into a database change, an email, a payment, or another automated process creates a different boundary: the model now has a path from information it reads to systems it can change. Better alignment still matters, but the surrounding workflow determines what the model is authorized to do.
For the past two years, much of the enterprise AI security discussion has centered on a narrower interaction: a user submits a prompt, a large language model returns text, and a human reads the response. That pattern puts alignment, jailbreaks, and hallucinations near the center of the security discussion. Production agents change those assumptions because they consume live data, invoke external tools and APIs, alter databases, launch downstream automation, and sometimes proceed without human review.
Those capabilities expand the security question from model behavior to the authority attached to model decisions. Teams have to ask what authority a probabilistic component has, which data can influence its intermediate decisions, and what happens when either the input or the decision is wrong. The central production risk is that conventional workflow design can turn manipulation or ordinary model error into an authorized action on a real system.
Untrusted inputs become dangerous when they inherit real authority
That risk begins before an agent calls a tool because connected data becomes part of its decision-making context. An agent might retrieve a CRM record, support ticket, scraped webpage, or shared document and use its contents to decide what to do next. An attacker can therefore influence behavior by changing material the model will eventually read, without direct access to the model.
That indirect route can be hard to recognize because hostile material can arrive inside ordinary business content. Instructions buried in a PDF, an email signature, or a product review can enter the same context the agent uses to interpret its task. Once there, they can redirect a later decision much as a malicious instruction entered directly into a prompt could. Connecting an agent to more information consequently expands the set of parties and documents that can influence its behavior.
That influence becomes consequential when the agent also has useful credentials. Agent frameworks can let a model send email, query a database, modify a record, or execute code, so the model’s judgment can determine which operation runs and with which arguments. A manipulated instruction then has a route from untrusted content to an authenticated tool.
That route can become more powerful during development because broad API keys and service accounts with extensive access are often the fastest way to get a demonstration working. In that setting, permissions can reflect “what’s convenient during development.” If those credentials survive the move into production, or a successful demo leads directly to a full-permission integration, the agent receives authority chosen for implementation speed rather than the production task.
The same credential problem applies when the model makes an error. A reasoning mistake combined with a broadly privileged credential can delete the wrong record, send a message to the wrong recipient, or trigger a payment. The credential authorizes the operation presented to it whether the choice was intended, hallucinated, or induced by hostile content.
Prompt injection and excessive permissions therefore compound each other. Injection determines how an attacker may influence the agent’s decision, while permissions determine how much that influenced decision can change. Reducing either side reduces the possible damage, but production security has to address the whole causal path: outside content enters context, changes a decision, and reaches a tool carrying legitimate authority.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.
Agent pipelines can carry untrusted influence across hidden trust boundaries
That causal path becomes harder to see when several models and tools participate in a workflow. Multi-tool and multi-agent pipelines routinely pass one component’s output to another, and a downstream component may accept that output without re-validating the material that shaped it. Processing can change how the data looks while preserving an outside party’s influence over it.
Consider a model that reads an untrusted document and produces a summary for a second agent. If the second agent has write access to a production system and treats the summary as trusted input, content controlled by an outside party can indirectly influence a production write. No human explicitly gave that party production access, yet the pipeline creates a decision path to a component that already has it.
That path persists because summarization does not inherently make untrusted material trustworthy. A transformation may shorten, restructure, or reinterpret content while preserving an embedded instruction’s effect on downstream behavior. Security decisions therefore have to follow provenance, the origin and processing history of data, and its influence across component boundaries, so an earlier model’s transformation does not become an implicit security guarantee.
Following provenance becomes especially important when authority increases downstream. The component exposed to outside content might have little power, while the next component can alter production data. Re-validation at those handoffs prevents a low-trust input path from quietly becoming a high-authority action path.
Autonomy multiplies errors, while agent execution can be harder to reconstruct
Once a pipeline can reach privileged actions, greater autonomy changes the scale of a failure. A workflow can move from drafting a response to sending it, or from identifying an anomaly to acting on it, removing an interruption that previously gave a person a chance to catch an error. Without a rate limit, approval gate, or kill switch between decision and execution, one faulty choice can start many automated actions.
That first action matters because the resulting problem can accumulate. A bad reasoning chain can keep making tool calls or triggering downstream processes until something else stops it, potentially producing “thousands of bad actions” before a person notices. Autonomy consequently changes the expected effect of an error because the workflow can repeat or extend the consequences without waiting for another human decision.
As those consequences accumulate, incident response can become harder. Conventional application security can use a logged request and call-stack tracing to reconstruct how software reached a result. Agent reasoning traces can vary non-deterministically, tool-call sequences can differ between runs, and teams may lack granular records of the model’s intermediate decisions.
That variability makes a basic incident question difficult: “What did the system actually do, and why?” An operator who can see only that “the model decided” and then “the action happened” lacks the evidence between those events. Re-running the same request may produce a different sequence, so reconstruction depends on preserving the original inputs, decisions, and calls as execution occurs.
Those execution risks do not require every workflow to operate manually. Autonomy has degrees: a model can perform substantial analysis and prepare an action while intervention requirements rise with the potential consequences. That design preserves useful model-driven workflows while reserving execution authority for stages where the possible impact justifies it.
Familiar production engineering belongs at the agent’s action boundary
Because the workflow grants execution authority, improvements to the model alone cannot secure the architecture around it. A probabilistic and occasionally unpredictable component consumes potentially untrusted material, makes intermediate decisions, chooses tools, and initiates actions, which makes ordinary integration weaknesses more consequential. Established production disciplines address those weaknesses through scoped permissions, input validation, observability, staged rollouts, and human review for consequential actions.
Least privilege starts by matching rights to the task. If an agent only needs to read customer records, its credential should permit reading and should exclude write access to those records. Removing unnecessary rights reduces the operations the workflow can execute when the model makes a mistake or hostile content influences it.
Once permissions are scoped, sandboxing controls how new integrations acquire production access. New tools and agent capabilities can first run in an environment where mistakes have low cost, followed by a deliberate, monitored move to production privileges. A successful demo shows that the workflow can perform its intended task; production permissions still require a separate decision based on the systems and operations that task needs.
After a capability reaches production, approval gates can contain actions with high consequences. A model can prepare a refund, draft an email to a customer, or propose a modification to a live system while explicit authorization remains necessary before the external action occurs. The possible blast radius provides a useful criterion for that gate because model confidence does not reduce the consequences of an erroneous execution.
Actions that pass those gates also need enough evidence for later investigation. Decision-level logging should record what an agent consumed, which tools it invoked, which arguments it supplied, and the rationale behind those decisions. Capturing that chain turns “what happened and why” into an incident-response question teams can investigate from the execution record.
Those records also need to preserve the origin of external content as it moves through the workflow. A webpage, uploaded file, or email can affect model behavior even when it arrives through a retrieval system rather than a chat box. Treating those inputs with the same suspicion conventional web applications give user input keeps their origin visible through processing and helps teams enforce controls before the resulting information reaches a component with stronger privileges.
Even with those preventive controls, circuit breakers limit damage after execution starts. Rate limits constrain how quickly automated actions accumulate, anomaly detection can identify unusual behavior, and manual override gives operators a way to stop execution. These controls matter because a single faulty reasoning chain can otherwise keep converting one initial failure into large numbers of erroneous actions before detection.
AI security ownership must expand beyond the model team
Putting those production controls around an agent requires ownership beyond model behavior. AI-security questions commonly land with ML or data-science teams, which can evaluate model behavior but may lack the organizational position or funding to own integration security, access control, and production observability. Once an agent gains write access to real systems, those surrounding responsibilities become part of its security whether or not they sit within the model team’s mandate.
Those responsibilities already overlap with the work of security and platform engineering teams, which have longstanding experience with service-to-service security, scoped credentials, audit logging, and incident response. The timing matters because these teams may enter only after an agentic workflow is live, when permissions and integration choices have already hardened into a production design. Organizations successfully scaling AI tend to close that organizational gap early and treat an agent as a production service once it can write to real systems.
That organizational gap directly affects how authority is granted. Ownership has to follow the authority given to the workflow, so model-mediated decisions that can change production systems need security and platform participation before those permissions are granted. Their existing controls then govern what those decisions are allowed to become in the systems the agent can change.
Key executive takeaways
- Control the action boundary: AI agents create a larger security risk once model decisions can trigger production actions. Security and platform teams should match controls to the authority each workflow receives.
- Limit the authority attached to untrusted inputs: Prompt injection and model errors become more damaging when agents hold broad credentials. Platform teams can reduce the blast radius by applying least privilege to every tool, API, and service account an agent uses.
- Preserve trust boundaries across agent pipelines: Summaries and other model transformations can carry untrusted influence into components with greater privileges. Engineering teams should track data provenance and revalidate inputs when authority increases downstream.
- Constrain autonomy and preserve execution evidence: Automated workflows can multiply a single faulty decision across many actions while non-deterministic execution complicates investigation. Operators should use approval gates, rate limits, kill switches, and decision-level logging for consequential workflows.
- Apply production controls to agent actions: Agent security depends on established controls including scoped permissions, sandboxing, staged rollouts, human approval, observability, and circuit breakers. Engineering teams should make these controls part of the path between model decisions and external actions.
- Assign ownership before granting production access: Agent security spans model behavior, identity, access control, integration security, and incident response. Security and platform teams should participate before an agent receives write access so production authority is governed from the start.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.


