AI expertise starts with production experience

Using ChatGPT, Claude, Gemini, or Copilot is now a basic business skill. It shows familiarity with generative AI. It says little about a person’s ability to deploy AI inside a large company.

Enterprise customer experience (CX) creates a much harder test. A production system must work with live customers, existing software, imperfect data, security controls, and real operating volumes. It must keep working when customer behavior changes. Employees must understand and trust its output. Someone must also own the financial and customer metrics when performance falls.

This creates four useful levels of AI experience. Consumers use AI assistants. Demonstrators build prototypes and proofs of concept. Implementers put AI into live operations. Operators run those systems over time, monitor performance, correct failures, retrain models when required, and remain accountable for results.

For executive hiring and procurement, the last two levels carry the strongest evidence. Ask candidates, consultants, and vendors what they put into production. Ask how many customers or transactions it handled. Ask what failed, how long integration took, what the full implementation cost, and which business metric changed.

A successful demonstration proves technical feasibility under controlled conditions. Production proves whether the company can turn that capability into reliable operations. That distinction matters because AI value depends on the entire operating system around the model: data, integrations, processes, controls, employee adoption, measurement, and ongoing maintenance.

The executive standard for AI expertise should therefore be measurable operating experience. A useful dashboard carries more information than a polished presentation. It shows volume, accuracy, containment, customer satisfaction, cost-to-serve, failures, and trends over time. These are the signals that reveal whether AI is creating durable business value.

Most enterprise AI initiatives fail to produce measurable business value

The failure numbers are severe. Research from MIT found that roughly 95% of generative AI pilots delivered no measurable impact on profit and loss, with about 5% producing meaningful revenue acceleration. RAND Corporation research found that more than 80% of AI projects fail, around twice the failure rate of non-AI technology projects.

The central constraint is execution. RAND identifies organizational causes including leaders misunderstanding the problem, inadequate data, and teams pursuing technology instead of a defined business outcome. These weaknesses can undermine a project even when the underlying AI model works as designed.

This changes how executives should define an AI project. Start with an economic or operational problem. Establish the baseline. Define the metric that AI must move. Determine what data and system changes are required. Then select the technology. A project aimed at reducing customer-service costs, for example, needs a baseline cost-to-serve and a clear method for measuring changes after deployment.

Pilot success also needs a tougher definition. A model can produce accurate answers during a demonstration while creating no measurable financial return. The real test comes when it connects to production systems, processes live workloads, operates within compliance requirements, and changes customer or employee behavior at sufficient scale.

The reported 80–95% failure range should influence portfolio management. Large AI budgets do not require large numbers of uncontrolled experiments. Executives can fund staged programs with explicit gates for data readiness, technical performance, adoption, operational reliability, and financial impact. Projects that fail those gates should be redesigned or stopped. Projects that demonstrate value can receive more capital.

The small successful group is equally important. MIT’s figure of roughly 5% achieving meaningful revenue acceleration shows that enterprise AI can generate business value. The practical question is therefore how to reproduce the conditions behind those successes.

That requires disciplined execution. Give each initiative a measurable outcome, a responsible owner, suitable data, integration resources, and a path from pilot to production. Track the result against the original baseline. AI investment becomes much easier to govern when every project must answer one question: what measurable business outcome changed?

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.

Data quality and integration determine whether enterprise AI works

Enterprise AI depends on the quality and accessibility of company data. This is often the main deployment constraint. A capable model still produces weak business results when customer records are duplicated, systems disagree, historical data is incomplete, or key information cannot move between applications fast enough.

Customer experience environments make this problem especially visible. One customer may appear differently across the CRM, billing platform, product database, and support system. Service records may contain vague resolution notes. Sensor feeds can have gaps. These problems reduce the context available to AI and weaken the accuracy of predictions and automated decisions.

Data preparation therefore belongs in the core AI investment plan. Teams need to identify authoritative data sources, resolve duplicate records, establish common definitions, address missing information, and define clear rules for data ownership and access. This work can consume months, but it determines whether later AI investment can scale.

Integration creates the next constraint. Production AI may need connections to CRM software, ticketing platforms, data warehouses, telephony, and other legacy systems. Real-time applications add further requirements. Data pipelines must process peak workloads reliably and deliver information quickly enough for the model’s output to remain useful.

Deployment also continues after the initial release. Customer behavior changes. Products change. Fraud patterns evolve. These shifts can cause model drift, where performance deteriorates because current conditions differ from the data used to build or configure the system. Organizations need continuous monitoring, defined performance thresholds, and processes for updating or retraining models.

Employee adoption has a direct impact on value as well. Analysts, service agents, and operations teams need enough confidence in an AI recommendation to act on it. That confidence comes from observable accuracy, clear escalation rules, and evidence that the system performs reliably in their workflows.

For executives, this changes the AI budget conversation. Model and software costs represent only part of the investment. Data cleanup, integration, process redesign, monitoring, retraining, and change management belong in the full cost calculation. Funding these capabilities early reduces the risk of building pilots that cannot survive production.

The right executive questions are concrete. Which systems contain the required data? Which system has the authoritative customer record? How much data needs correction? What must operate in real time? Who owns data quality? How will model performance be monitored after deployment? Clear answers indicate whether an AI program is ready to scale.

Chatbots capture only a narrow share of AI’s value in customer experience

Chatbots have become one of the most visible uses of generative AI in customer experience. They can answer routine questions, automate common requests, and reduce some service workload. Their visibility can also distort investment priorities when leaders treat conversational automation as the main AI opportunity.

Many higher-value applications operate earlier in the customer journey. AI can detect fraud during a transaction, identify a manufacturing defect before delivery, predict equipment problems that could disrupt service, detect billing anomalies before customers complain, identify customers at risk of leaving, and adjust offers or experiences using current behavior.

McKinsey’s 2026 analysis provides concrete operational examples. JPMorgan uses AI to scan transactions for fraud in real time. BMW applies computer vision, which allows software to interpret images, to detect manufacturing defects on production lines. Siemens combines predictive maintenance with production planning to reduce downtime. American Express adjusts offers using transaction behavior in near real time.

These systems affect customer experience without requiring a customer to interact directly with AI. Preventing a fraudulent payment or defective shipment can remove the need for a later service interaction altogether. This expands CX strategy from handling customer contacts to identifying and preventing the events that create those contacts.

Other useful applications sit directly inside customer operations. Voice-of-customer systems can analyze large volumes of calls and identify recurring issues. Churn models can detect early signals that a customer may leave. Real-time decision systems can adjust journeys while an interaction is underway. Anomaly detection can identify unusual account or billing behavior that needs investigation.

This broader view changes how executives should prioritize investment. Start with the customer outcome and trace the operational process that creates it. Identify where better prediction, classification, detection, or automated decision-making can remove cost, delay, errors, or customer effort. Then assess whether AI is technically and economically suitable.

Chatbots remain useful where conversational automation solves a clearly measured problem. They should compete for capital against fraud detection, predictive service, personalization, quality control, churn prevention, and other AI applications on the same basis: measurable customer and financial outcomes.

For C-suite leaders, the strategic opportunity is therefore larger than customer-service automation. AI can influence product quality, reliability, risk, personalization, and problem prevention. The strongest AI portfolio will place investment where those capabilities produce the greatest measurable improvement in customer experience and business performance.

Advanced AI applications show how much value sits inside core operations

Some of the strongest AI applications work inside business processes that directly shape the customer experience. Fraud detection, manufacturing quality, predictive maintenance, personalization, and scientific discovery demonstrate what becomes possible when AI has access to relevant data and is integrated into operational decisions.

McKinsey’s 2026 analysis highlights several examples. JPMorgan uses AI to scan transactions for potential fraud in real time. BMW applies computer vision to identify defects during manufacturing. Siemens combines predictive maintenance with production planning to reduce downtime. American Express uses transaction behavior to adjust customer offers in near real time.

Each case connects AI to a specific operational outcome. Fraud detection protects customers and reduces losses. Manufacturing inspection can prevent defective products from reaching buyers. Predictive maintenance can improve reliability and reduce delays. Real-time personalization can make offers more relevant to an individual customer’s current behavior.

The same principle applies to CX investment. Executives should identify the events that create customer cost, effort, dissatisfaction, or risk and determine whether AI can detect or predict those events earlier. This approach expands the investment opportunity across operations, risk, manufacturing, product, and service.

Scientific AI shows the potential at much greater scale. AlphaFold has predicted the structures of about 200 million proteins and has been used by more than two million researchers across 190 countries. Protein structure is important because a protein’s three-dimensional form helps scientists understand its function and interactions, making structure prediction valuable in biological research and drug discovery.

AlphaFold also shows what happens when AI is applied to a well-defined problem with large amounts of relevant data and a measurable output. For executives, the broader lesson is practical: the value of AI depends heavily on selecting a problem where prediction or automated analysis can materially improve an existing process.

This requires a portfolio view of AI. Customer-facing conversational systems should compete for investment alongside fraud prevention, quality control, predictive operations, churn management, personalization, and other applications. Capital should flow toward the use cases with the clearest combination of customer impact, operational feasibility, and economic return.

Customer behavior and ROI data put company-built chatbots under pressure

The economics of customer-service AI deserve close scrutiny. A 2026 Gartner survey found that customers were three times more likely to use a general-purpose AI tool than a chatbot built by a company. Usage of company chatbots had also remained flat since 2022.

The investment numbers make this more important. Gartner found that service leaders allocated a median 12% of their budgets to AI, described as the highest share among business functions. Yet only 24% of service leaders could demonstrate positive returns.

Together, these figures identify a significant management problem. Companies are directing substantial resources toward AI while many service organizations still cannot connect those investments to positive financial results. Customer behavior also indicates strong demand for general-purpose AI products.

There are practical reasons executives should examine this gap carefully. General-purpose AI services can give users a broad interface for research, comparison, explanation, and problem-solving. A company chatbot usually operates within a narrower set of company systems and approved workflows. Its value therefore depends heavily on whether it has accurate customer context, reliable access to enterprise data, and enough authority to complete the customer’s task.

This raises the standard for a company-built assistant. Answering questions is useful, but executives should measure whether the system resolves customer needs. Relevant measures include successful resolution, containment of requests that do not require human assistance, customer satisfaction by intent, repeat-contact rates, escalation rates, cost-to-serve, and completion of key transactions.

A chatbot can still justify investment when those measures improve. The decision should come from observed customer behavior and economics. A high adoption figure has limited value when customers repeatedly escalate to human agents, return with the same problem, or leave without completing the intended task.

Executives should therefore examine AI service spending at the use-case level. Ask how much each implementation costs to build, integrate, operate, monitor, and update. Compare those costs with measurable changes in service demand, customer outcomes, revenue, and operating expense. This gives leadership a clearer basis for deciding which systems to expand, redesign, or stop.

The 12% budget figure and 24% positive-return figure should serve as a governance signal. AI spending can grow faster than demonstrated value. Strong governance links each investment to an accountable owner, a baseline, defined customer and financial metrics, and a timetable for proving results. That discipline allows companies to keep investing in AI while directing capital toward applications customers actually use and outcomes the business can measure.

Five signals reveal whether AI expertise is grounded in production

AI credentials are easy to claim. Production experience is easier to verify when executives ask questions tied to actual deployments. Five signals provide a practical test: specificity, failure fluency, boundary awareness, cost honesty, and ownership.

Specificity comes first. Experienced operators can describe scale and results with precise numbers. A statement such as “34% of tier-one tickets out of 60,000 per month” gives leaders a measurable result and its denominator. Claims such as “transformative” or “seamless” provide little information for assessing performance. Executives should ask for volumes, time periods, baselines, and measurement methods behind every major percentage.

Failure fluency is equally important. Production systems encounter bad data, integration failures, unexpected customer behavior, latency problems, inaccurate outputs, and operational resistance. Someone who has managed a deployment should be able to explain what failed, why it failed, how the team detected it, and what changed afterward. These answers often reveal more operational knowledge than a successful demonstration.

Boundary awareness shows whether someone understands where automation creates unacceptable risk. Experienced practitioners can identify decisions that require human review, situations where confidence is too low for automated action, and customer interactions where escalation rules are necessary. Those boundaries should reflect the financial, regulatory, reputational, and customer consequences of an error.

Cost honesty exposes whether an AI business case reflects production economics. Software licenses and model usage are only part of total cost. Data preparation, integration, testing, security, monitoring, retraining, process redesign, employee adoption, and ongoing operations can materially change the return. Executives should require total implementation and operating costs before approving ROI claims.

Ownership completes the test. Credible operators know which metric they were accountable for and what happened when it deteriorated. In CX, that could include containment, customer satisfaction by intent, model accuracy, repeat contacts, or cost-to-serve. Accountability forces teams to connect model behavior with operational and financial consequences.

These five signals can strengthen hiring, procurement, consulting engagements, and investment reviews. Ask candidates and suppliers what they deployed, at what scale, what broke, what they refused to automate, what the full deployment cost, and which metric they personally owned. Specific answers with evidence indicate production experience.

For C-suite leaders, this approach also improves governance. AI competence becomes something the organization can test through operating evidence. That makes it easier to identify teams capable of moving from experimentation to reliable production systems.

Theory creates value when it is built from verifiable operating evidence

Companies need both practitioners and people who can turn practical experience into reusable knowledge. Practitioners discover what works through deployment. Researchers, strategists, consultants, and other thought leaders can organize those lessons into methods that other teams can apply.

The sequence matters. Practice should generate the evidence first. Theory can then explain the conditions behind the result, identify patterns across implementations, and determine which lessons may transfer to other organizations.

This approach addresses an important scaling problem. A deployment team may learn how to clean a difficult dataset, integrate AI with legacy systems, establish human escalation rules, or monitor model drift. Those lessons often remain inside a business unit or company. Careful synthesis can convert specific implementation experience into guidance that saves other teams time and money.

Evidence determines whether that synthesis is useful. Strong AI guidance should connect claims to identifiable cases, measurable outcomes, deployment conditions, and credible data. It should also distinguish observed results from forecasts and assumptions. Executives can apply a simple “footnotes test”: trace important claims back to the evidence supporting them.

This discipline matters because results from one company do not automatically transfer to another. AI performance depends on data quality, process design, system architecture, customer behavior, regulation, operating scale, and organizational capability. A method built from several real deployments can identify which conditions are essential and which practices are specific to one environment.

Practitioners also benefit from strong theoretical work. Individual teams see a limited number of implementations. Researchers and experienced synthesizers can compare results across companies and industries, identify recurring failure patterns, and provide frameworks for evaluating new use cases. Done well, this reduces repeated experimentation and improves capital allocation.

For executives, the goal is a productive relationship between implementation and synthesis. Give operating teams the authority to test ideas against real outcomes. Require strategy teams and external advisers to support recommendations with traceable cases and measurable evidence. Separate proven findings from assumptions in investment documents.

This creates a stronger learning system for enterprise AI. Production generates evidence. Careful analysis turns that evidence into reusable knowledge. New deployments then test and refine the resulting methods. Over time, that cycle can improve execution quality while giving leadership a firmer basis for major AI decisions.

AI investment needs a higher burden of proof

Executives should judge AI-driven CX programs by production evidence. Every major initiative should answer four questions: What shipped? What did it cost? What broke? What measurable outcome changed?

These questions shift attention toward execution. A production deployment must operate with live customers, real transaction volumes, existing systems, imperfect data, security requirements, and changing behavior. Its performance must remain measurable after the initial launch.

Financial accountability is essential. AI can create business value when the use case is clearly scoped, the required data is reliable, and the deployment has an accountable owner. The business case should define the baseline before implementation and track the resulting change in revenue, cost-to-serve, retention, fraud losses, customer satisfaction, resolution rates, or another relevant metric.

Total cost also needs careful treatment. Executives should include data cleanup, integration, infrastructure, security, testing, monitoring, retraining, employee adoption, and ongoing operations. These costs determine the economic result of a production system and should appear in investment decisions from the beginning.

Failure evidence has value as well. Teams with production experience should be able to explain where models performed poorly, which integrations failed, what customers did unexpectedly, where human intervention became necessary, and how those problems affected business metrics. This information helps leadership understand operational risk and determine whether teams can manage it.

Dashboards provide the strongest basis for ongoing governance. In CX, useful measures can include containment, resolution rates, customer satisfaction by intent, repeat contacts, model accuracy, cost-to-serve, and escalation rates. Financial measures should connect operational improvements to the P&L wherever possible.

Research on AI project performance reinforces the need for this discipline. MIT research in the material found that roughly 95% of generative AI pilots produced no measurable P&L impact, while about 5% achieved meaningful revenue acceleration. RAND Corporation research found that more than 80% of AI projects fail, roughly twice the failure rate of non-AI technology projects. RAND linked these failures to issues including poorly understood problems, inadequate data, and technology-led initiatives without sufficiently clear outcomes.

These figures support a selective approach to investment. Executives can fund experimentation while requiring evidence at each stage. A project should progress from prototype to pilot to scaled deployment only when it meets predefined thresholds for data readiness, technical performance, operational reliability, customer impact, and economics. Weak projects should be corrected or stopped before they absorb larger budgets.

The same burden of proof should apply to vendors, consultants, candidates, and internal AI leaders. Request deployment volumes, baselines, implementation costs, failure histories, outcome metrics, and clear ownership. Specific evidence makes expertise easier to evaluate.

AI raises the standard for CX leadership because successful deployment requires technical capability, operational discipline, and financial accountability at the same time. The strongest investment case is therefore a measurable one: a defined customer problem, a working production system, a responsible owner, a complete cost base, and verified improvement in business outcomes.

The bottom line

AI in customer experience has moved beyond the stage where access to powerful models creates an advantage. The advantage now comes from execution. Clean data, reliable integration, production discipline, clear ownership, and measurable economics determine which companies create value.

For executives, this should change the questions asked in every AI review. What is running in production? At what scale? What broke? What does it cost end to end? Which customer and financial metrics changed? Who owns those metrics when performance drops?

The same standard should guide hiring, vendor selection, and capital allocation. Look for practitioners who can discuss failures as precisely as successes. Require denominators behind percentage improvements. Include data preparation, integration, monitoring, and retraining in the economics. Demand evidence that survives beyond the pilot.

AI can produce substantial value in CX. The strongest opportunities may sit in fraud prevention, quality control, predictive service, churn management, personalization, and problem prevention as much as conversational interfaces. Start with the customer outcome and operating constraint, then determine where AI can materially improve it.

The next phase of enterprise AI will reward companies that make deployment evidence their standard for credibility. Strategy sets direction. Production proves whether it works. For the C-suite, that proof should decide where the next dollar goes.

Alexander Procter

August 24, 2026

18 Min

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.