AI can pass the gate and still fail the customer

A passing AI evaluation is weak evidence of what a customer will actually experience in production. In July, 49% of 108 people surveyed by VB Pulse at companies with at least 100 employees said their organization had experienced a customer-visible problem after an AI agent or LLM-powered feature cleared internal testing. The June figure was 50%, and 24% of July respondents said this had happened more than once.

What the 49% figure measures defines how far that result can go. It captures organizations that experienced at least one qualifying incident during the previous year; it does not measure the share of individual agent runs that failed or the failure rate of an evaluation product. Because a company deploying many agents has more opportunities to experience an incident, production exposure can affect whether it enters that group.

Within that boundary, the result still limits how much responsibility enterprises can place on a release gate. Internal evaluation tests behavior against criteria chosen in advance and can help decide whether an AI system proceeds toward deployment, but roughly half of surveyed organizations in both waves had seen a system clear such testing and later create a customer-visible problem. A passing evaluation by itself cannot establish production reliability.

Confidence is rising while the reported miss rate is flat

That limit makes the change in confidence more striking. In July, 13% of respondents expressed complete trust in automated evaluation, up from 5% in June, while the share naming poor alignment between tests and real-world results as their largest concern fell from 29% to 19%. The incident measure moved by only one percentage point over the same period.

Measure June July
Complete trust in automated evaluation 5% 13%
Poor test/real-world alignment named as largest concern 29% 19%
Organization had at least one test-passing, customer-visible incident 50% 49%

Those movements show better sentiment toward automated evaluation without a comparable change in the reach of this particular customer-visible outcome. Because the incident measure records organizational experience, it cannot show how often failures occurred within each organization. It does show that the proportion of surveyed companies reporting at least one such miss was effectively flat.

The flat incident measure extends what VentureBeat’s June research called the “enterprise evaluation gap.” In June, agent authority was increasing faster than enterprises’ ability to test agents dependably. July adds a different observation: organizations can report greater confidence in automated evaluation while the share reporting at least one test-approved system that later caused a customer problem changes little.

The survey design keeps that comparison narrow because organizational experience differs from a controlled test of evaluation quality. More deployments create more chances to register at least one incident, and changes in the systems companies deploy can alter that exposure. The July results establish a divergence between reported confidence and the prevalence of organizations reporting this type of miss; they do not establish its cause.

Other month-over-month movements add context to the changing buying environment. VentureBeat Intelligence identified four developments: complete confidence in automated evaluation increased, concern about poor real-world alignment decreased, more respondents chose integration ease as the decisive buying factor, and Braintrust gained primary-platform share. The first two track the confidence shift above, while integration ease and Braintrust show that evaluation purchasing preferences were changing at the same time.

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.

Companies that experienced eval misses are more skeptical and more automated

The July cross-tab sharpens the divergence because firsthand experience is associated with much lower confidence. Among enterprises where a tested AI feature later disappointed a customer, only 4% placed complete faith in automated checks. Among organizations with no identified comparable incident, 24% did so, a sixfold difference.

The counts behind those percentages show the size of the groups being compared. Two of 53 previously burned enterprises expressed complete faith in automated checks, compared with 10 of 41 enterprises with no identified testing miss. The association is clear: respondents that had seen the release gate miss a customer-facing problem were much less likely to place complete confidence in automation.

Lower confidence, however, did not correspond to less interest in autonomous deployment. Across the July sample, 67% either already allowed an agent to push code or change a system without human approval in certain low-risk cases, or were changing their pipelines to allow that during the coming year. The overall proportion was unchanged from June, with 37% already permitting the practice in limited cases and another 30% building toward it.

Within that movement, the burned group was pursuing the model more aggressively. Among enterprises that had experienced a test-approved system disappointing a customer, 85% were pursuing deployment without human approval in the relevant cases, compared with 61% among enterprises reporting no comparable incident. Experience of an evaluation miss is therefore associated with lower confidence in automated checks and greater pursuit of deployment autonomy.

The same relationship appears among respondents rejecting full deployment automation for the years ahead. Only 11% of burned respondents rejected end-to-end automation, compared with 24% of unburned respondents. A customer-visible failure, in this sample, was therefore associated with less rejection of end-to-end automation.

That association makes the distinction between human approval and automated evaluation important. An organization can lose confidence that an automated check captures every production failure while still concluding that requiring a person to approve every low-risk production change is an unsuitable operating model. The July results leave the reasons for that judgment unresolved, but they show that firsthand misses have not stopped movement toward greater autonomy.

Correlation leaves the cause unresolved

The stronger move toward autonomy among burned companies does not show that the failure caused the move. The 85% relationship could reflect greater risk appetite or other differences between the groups. Deployment maturity could also contribute because organizations operating more agents at greater volume have more opportunities to encounter an incident, while those same organizations may have engineering infrastructure that makes automated deployment easier to implement.

The subgroup sizes further limit conclusions about cause. Burned and unburned comparisons use groups ranging from 41 to 53 respondents, while other cross-tabs use groups from 40 to 68. VB Pulse was self-selected rather than a probability sample, so the results provide directional evidence but cannot simply be projected across enterprises generally.

Month-to-month comparisons face another constraint because the July respondent mix differed from June. Technology and software participation fell nine percentage points to 14%, while retail and consumer participation rose four points to 19%. Changes between the waves can therefore reflect a changing sample alongside shifts in enterprise attitudes or practices.

The full sample shows the scale and decision context of those limits. The two VB Pulse waves from VentureBeat Intelligence comprise 265 enterprise responses: 157 in June and 108 in July. In the July sample, 63% worked at organizations with 100 to 2,499 employees, and 69% described themselves as final AI-buying authorities or people who recommend or influence AI purchases. Those respondents make the findings relevant to enterprise technology decisions, while self-selection and the changing industry mix constrain wider market conclusions.

Within those boundaries, the association still matters for engineering leaders because experiencing a customer-visible incident does not appear to halt movement toward deployment autonomy. The data leave recklessness, maturity and other possible explanations for that association unresolved. Instead, they show that evaluation skepticism and automated deployment can coexist inside the same organizations.

Complex agents make exhaustive pre-production evals harder to maintain

That coexistence raises a practical question: what controls can support more autonomy when pre-production tests cannot anticipate every behavior? Raindrop.ai, an automated agent error-monitoring and mitigation platform, says it is seeing companies change how they approach evaluation, according to CTO Ben Hylak. Raindrop.ai sells monitoring and mitigation for agent errors, so it has a commercial interest in enterprises placing more value on continuous detection. In a direct message to VentureBeat, Hylak said, “We are seeing the great-decline of evals as we know them,” describing a shift he said had developed from “just a few months ago.”

Hylak links that claimed shift to large enterprises and increasing system complexity. “The Fortune 100 are increasingly reducing eval sets and deprioritizing maintenance. As systems grow more complex (MCPs, subagents, etc.) it becomes impossible to fully enumerate the failure cases. Instead, they’re leaning on anomaly and issue detection solutions, both before and after production.” Hylak presents this as a pattern Raindrop.ai says it observes across the cohort.

The complexity Hylak describes comes from the range of behaviors that interacting components can produce. MCPs, or Model Context Protocol integrations, can connect agent systems to external tools and data, while subagents divide work among additional agents. As those components interact, maintaining a finite evaluation set, a bounded collection of predefined tests, that covers every relevant failure case becomes progressively harder because the organization must anticipate combinations of behavior before deployment.

That difficulty changes the purpose of anomaly and issue detection. A finite eval set asks whether known tests pass at a defined point before release, while continuous detection looks for unexpected or problematic behavior around production, including cases designers did not enumerate in advance. In Hylak’s account, Fortune 100 companies are shifting maintenance effort toward those detection mechanisms both before and after production.

Hylak’s vendor observation should remain distinct from the VB Pulse findings because the two evidence streams answer different questions. The survey establishes enterprise attitudes and reported incidents rather than changes in eval-set size, adoption of Raindrop.ai, or reasons for moving toward anomaly detection. Hylak’s account likewise describes a monitoring pattern his company says it sees as agent systems become harder to enumerate exhaustively, rather than a causal explanation for the survey’s changing confidence numbers.

More autonomous deployment expands the safety boundary

The move toward autonomous deployment makes controls outside a single release decision more important. If deployment volume rises while the probability of failure per deployment remains constant, the absolute number of incidents could rise even if the percentage of companies experiencing at least one incident stays flat. Because July measures whether an organization experienced a qualifying incident rather than total incident volume, that increase is an implied risk rather than an observed result.

That risk broadens the operational response beyond the approval decision. Human approval can be one control, while detection before and after release can catch problems as automated systems make more changes. Organizations increasing deployment autonomy therefore need ways to identify behavior that pre-production tests did not anticipate and respond when it appears in production.

The broader operating model also changes what “safety” has to cover. Evals remain useful for known behaviors and release criteria, but a passing score carries limited information about conditions the tests did not capture. As deployment becomes more autonomous, the reliability burden extends across the lifecycle of the deployed system, with detection operating on both sides of production.

Key takeaways for leaders

  • Treat passing evals as limited evidence: About half of surveyed enterprises reported a customer-visible problem after an AI system cleared internal testing. AI owners can use pre-production evals for known behaviors while maintaining controls for failures those tests do not cover.
  • Reconcile rising confidence with persistent misses: Complete trust in automated evaluation rose from 5% to 13%, while the share reporting at least one customer-visible miss stayed essentially flat at 49%. Technology leaders can track production outcomes alongside eval scores before expanding reliance on automated gates.
  • Prepare autonomy for imperfect evals: Among companies that experienced an eval miss, 85% were pursuing deployment without human approval in relevant cases, versus 61% of companies without a reported miss. Engineering organizations increasing autonomy need detection and response mechanisms that operate after release.
  • Avoid drawing causal conclusions from the survey: The link between prior AI failures and greater deployment autonomy could reflect deployment scale, maturity, risk tolerance or other factors. Decision-makers should treat the findings as directional because the survey was self-selected and subgroup sizes were small.
  • Extend evaluation into continuous detection: Complex systems using tools, MCP integrations and subagents create combinations of behavior that finite eval sets become harder to enumerate. Platform teams can pair defined pre-production tests with anomaly and issue detection before and after deployment.
  • Expand the safety boundary with deployment autonomy: Higher automated deployment volume creates more opportunities for incidents even when the failure rate per deployment does not change. Organizations removing human approvals from low-risk changes need lifecycle controls that detect unexpected behavior and support rapid production response.

Alexander Procter

October 8, 2026

10 Min

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.