A vendor that passed review may no longer be the vendor you assessed

A third-party vendor can pass a review and later add capabilities that change the basis on which the product was approved. Gartner calls one version of this problem AI creep: vendors add AI capabilities to existing products while customers may not fully understand what changed. Agentic capabilities raise the stakes because they allow AI to take actions. As vendors add those capabilities quickly, an assessment can describe an earlier version of the product rather than the one an organization now uses.

That gap was a key issue at Gartner’s 2026 Enterprise Risk, Audit & Compliance Conference, where multiple sessions examined how added AI capabilities expand an organization’s AI risk surface. The AI risk surface is the set of AI functions that can create security, compliance, operational or governance risk. In interviews with TechTarget, Gartner analysts Devanshu Mehrotra and Kjell Carlsson described how changing capabilities challenge established third-party assurance practices. Gartner sells technology research and advisory services, including guidance on risk and governance, so it has a commercial interest in organizations seeking advice on these problems.

The challenge starts with the lack of a single event that always triggers reassessment. A vendor might add a chatbot, improve search, introduce an agent or gradually expand what an existing agent can do. Each release can look incremental even as the product gains materially different authority and behavior from the version IT and risk teams approved. Assurance needs a way to detect relevant changes between formal assessment cycles.

AI creep breaks the assumptions behind point-in-time assurance

Detecting those changes puts pressure on three assumptions behind traditional third-party assurance, according to Devanshu Mehrotra, Gartner senior director and analyst. Vendor capabilities may not change slowly enough for periodic reviews to remain representative, contractual disclosures may not surface significant changes, and vendor attestations may not accurately describe what happens inside the product. Agentic AI can weaken all three because functionality and behavior may change after assurance work is complete.

The first pressure is on the useful life of an assessment because a product can gain new data access or authority after review. The findings still describe the state that teams assessed, but the operating product has moved beyond it. Mehrotra described the timing problem in his TechTarget interview: “Sometimes capabilities move faster than the vendor can keep up.” Customers can therefore face a gap between current product behavior and what their assurance process has formally evaluated.

The timing problem also affects disclosure because significant change can accumulate across several releases. A large, clearly announced addition can trigger legal, security or risk review, while smaller changes can each fall below an organization’s threshold for action even when their combined effect changes data use, authority or consequences. Materiality therefore has to account for accumulated change. Without that rule, teams can approve each step while missing what the steps create together.

That cumulative effect connects technical change to contract design. Mehrotra said, “Have a clear list of what is contractually allowed and not allowed, and a definition of what a material change is, because a bunch of incremental changes can add up to a material change.” A defined material-change threshold gives procurement, IT and risk teams a shared basis for deciding when successive AI additions require renewed scrutiny. The decision then depends on their combined effect rather than the apparent importance of one release.

Point-in-time artifacts have a related limitation because vendor attestations, SOC reports, questionnaires and contracts describe or govern a product at a particular stage. They remain useful inputs to third-party risk decisions, but a later AI capability can make some earlier conclusions stale. Teams can keep those controls while adding mechanisms to detect changes, obtain disclosure and collect evidence about current behavior. Completing the original artifacts cannot by itself establish that the capability set will remain stable until the next assessment.

Once those artifacts become a baseline, assurance can ask a more useful recurring question: what material change has occurred since they were produced? Answering it requires visibility into new AI functionality, criteria for deciding when another review is warranted, and evidence of what vendor systems do in operation. Periodic assessments can remain scheduled checkpoints while oversight continues between them. Continuous awareness does not require literal real-time surveillance of every vendor system.

Highly regulated industries already use this governance principle. Mehrotra pointed to pharmaceuticals, where initial approval is followed by monitoring of performance, evaluation of new information and scrutiny of significant changes. Applied to AI governance, that model preserves the initial decision while allowing later evidence to change what decision-makers need to evaluate.

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.

Risk follows the AI capability’s access and authority

Once assurance follows change over time, teams need a more precise unit of assessment. A single application can contain several AI capabilities with sharply different risks because their access and authority differ. Kjell Carlsson, Gartner vice president analyst, argues that CIOs should examine what AI features can actually do inside third-party products. Existing application classifications alone cannot describe the risk created by every capability within them.

Autonomous agents make the difference concrete. Mehrotra gave the example of an agent that provisions user access: if its authority or behavior is wrong, it could grant someone inappropriate access to sensitive information. An agent that summarizes emails each morning has a much smaller potential consequence. Both are AI agents inside enterprise software, but the actions they may perform and the resources affected by those actions create different risk.

That difference makes access, authority and consequence central to capability-level assessment. Teams need to determine which company data an AI capability can reach, which operations it may perform, and what happens when those operations are incorrect or inappropriate. In the access-provisioning example, the AI can alter another person’s permissions; in the morning-email example, its primary function is producing a summary. The capability’s behavior provides information that an application name cannot.

Behavior can still be hard to see because AI creep does not always arrive through a conspicuous autonomous agent. Carlsson identifies chatbots and agent-building tools as visible examples, while AI-powered search and anomaly detection can introduce AI through less obvious functions. A governance process focused only on prominent AI interfaces can therefore miss changes in how company data is processed or how decisions are supported. Visibility has to reach AI functions embedded deeper in the product.

For CIOs, engineering leaders and risk teams, that visibility changes the practical question during vendor review. Knowing that a product “uses AI” says little about a particular capability’s authority. Teams need to establish what the capability can access, what it can do with that access, who is accountable for it, and what organizational consequence follows when its behavior changes. Those answers provide the basis for deciding which vendor changes require further assurance.

Turn assurance into an ongoing evidence and disclosure process

Capability-level assessment makes continuing oversight more specific because teams can watch for changes in access and authority. IT and risk teams can track the AI capabilities a vendor adds, what each can do and how an expansion changes organizational risk. If an agent gains access to another data set or gains authority to perform another action, the new capability becomes an assurance event. Teams can then evaluate its significance against the state they previously approved.

Evaluating that event requires operational evidence of what happens in vendor systems. Mehrotra recommends seeking such evidence alongside point-in-time assurance so teams can examine system behavior after an assessment is complete. SOC reports, questionnaires and attestations provide formal evidence, while operational information helps show whether current behavior still fits the assumptions and permissions behind the approval. The two forms of evidence cover different points in the product’s life.

Contracts can turn those observations into enforceable expectations. Vendor requirements can specify when material AI-capability changes must be disclosed, define which behaviors are contractually permitted or prohibited, and establish what counts as a material change. That definition also needs to account for cumulative change because several incremental additions can eventually create enough access, autonomy or consequence to warrant fresh review. Contract language then supplies the threshold against which teams can judge new operational evidence.

Carlsson’s questions give CIOs a recurring framework for making that judgment. They can ask the same questions when a capability first appears and again when its data access, authority, purpose or vendor governance changes:

  • “Where is company data going?”
  • “Who is involved and accountable?”
  • “What can the AI access and do?”
  • “Why is the capability valuable?”
  • “When will the vendor notify customers about material changes?”
  • “How does the organization know the AI remains trustworthy and governed?”

Those questions separate the elements of the decision so teams can see what changed. Data location and access establish exposure, accountability identifies who owns decisions and consequences, and value establishes why the organization accepts the capability’s risk. Notification terms determine how customers learn about meaningful vendor-side changes, while the final question calls for continuing evidence that the AI remains trustworthy and governed. Revisiting each element keeps the review tied to current capability behavior.

Repeated use also keeps a new questionnaire from becoming another fixed snapshot. A capability can have an acceptable answer when introduced and a different answer after the vendor expands its data access or allows autonomous action. Continuing oversight detects those changed answers and tests them against the organization’s contractual and risk thresholds. The pharmaceutical practice Mehrotra cited offers the earlier analogy: initial approval remains useful while later performance, information and significant changes continue to receive scrutiny.

Keep a live inventory of capabilities and accountability

Repeated review creates an internal information problem because the number and behavior of AI capabilities can change across an organization’s software environment. Valence Howden, advisory fellow at Info-Tech Research Group, says tracking what AI agents can actually do across that environment can be difficult. He recommends maintaining a registry of the AI tools and agents the organization uses, what those systems do and where accountability resides. Info-Tech Research Group sells technology research and advisory services, so it also has a commercial interest in organizations adopting technology-management and governance practices.

A registry gives that recurring review a durable internal record. When a supplier changes an AI function, the organization can compare the new behavior with the previously understood state, examine how its authority changed and identify the accountable owner. Security, engineering, IT and risk teams can also use the record to connect a vendor-side change with affected internal systems, data and decisions. Keeping the registry current makes it useful between individual vendor conversations.

That record can hold the distinct information raised by the Gartner and Info-Tech Research Group analysts without treating their recommendations as equivalent. Mehrotra’s emphasis on continuing visibility and operational evidence identifies changes that may have moved a product beyond an earlier assessment, while Carlsson’s recurring questions identify information teams can reassess as a capability evolves. Howden’s registry provides a place to retain capability, behavior and ownership information as the software environment changes. The resulting assurance record is useful only while it reflects the vendor capabilities the organization currently uses.

Main highlights

  • Reassess vendors as AI capabilities change: CIOs and risk teams need ongoing visibility into AI additions because new access, autonomy or data use can make an earlier vendor assessment stale. Define when individual or cumulative changes trigger renewed scrutiny.
  • Assess access and authority at the capability level: Risk teams should evaluate what each AI capability can access, which actions it can perform and the consequences of errors. Agentic functions with authority over permissions, data or operations warrant greater scrutiny.
  • Require evidence and disclosure between reviews: Procurement and risk teams should define material AI changes in contracts, require vendor notification and collect operational evidence of current behavior. These controls help identify when evolving capabilities exceed previously approved boundaries.
  • Maintain a live AI capability inventory: IT and governance teams should record vendor AI tools and agents, their functions, access and accountable owners. Keeping that registry current helps connect vendor changes to affected systems, data and business decisions.

Alexander Procter

October 5, 2026

10 Min

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.