Cheaper AI tokens are sending a misleading signal to enterprise budgets. Sachin Chitturu, partner and Southeast Asia leader at QuantumBlack, AI by McKinsey, told Tech Week Singapore that Silicon Data figures show per-token prices have fallen about 90% since 2023. Yet Gartner estimates cited by Chitturu put token consumption for agentic models, which can pursue multi-step tasks with greater autonomy, at five to 30 times more per task. Falling unit prices can therefore coexist with rising total bills.

“The cost of tokens is not reducing as much as the usage of tokens is increasing,” Chitturu said. That relationship changes the architecture question for technology leaders because model economics depend on both the price of each token and how many tokens a workload consumes. As agents take on longer tasks, enterprises must keep deciding which work deserves expensive reasoning, which can run on cheaper models, and where people should remain involved. QuantumBlack and McKinsey advise enterprise clients on AI, so they have a commercial stake in demand for the architecture and transformation work Chitturu discusses.

Cheaper tokens can still produce higher enterprise AI costs

Rising consumption weakens a common assumption in enterprise AI planning: improving model economics will steadily make deployed workloads cheaper. Agentic systems can consume the unit-price savings by doing more work within each task, so the relevant cost is the complete workload. For a CTO or FinOps team, financial control consequently depends on workload design, model choice and deployment mode together.

Those variables also change what standardisation means. Choosing one leading model provider may simplify an initial rollout, but each expanding workload then remains exposed to that provider’s economics and the model’s token consumption. Because simpler work and difficult reasoning have very different requirements, the useful architecture question becomes where each request should run and how its route should change as models and prices change.

The cost squeeze is already constraining adoption and exposing the ROI gap

The need for workload-level control is already visible in enterprise budgets. Chitturu cited an enterprise AI financial operations, or FinOps, survey in which 93% of enterprises reportedly overspent their AI budgets during the previous six months. FinOps applies financial management practices to technology consumption, and AI makes that discipline harder because applications, agents and models can change compute consumption quickly. Budget control must consequently account for how usage evolves after deployment.

That pressure is already restricting demand. In McKinsey’s latest global state-of-AI survey, one in five respondents said operating costs, including token costs, had caused their organisation to curb AI use. Technology organisations reported the highest rate of curtailment at 25%, followed by financial institutions at 24%. Even among top-performing organisations, Chitturu said fewer than 30% of employees report having access to all the AI capabilities they want.

Those restrictions matter more when individual productivity is compared with company-level financial returns. McKinsey found that 80% of respondents using AI at work said it improved their individual productivity, while 37% reported a positive contribution to organisational earnings before interest and taxes, or Ebit. Just 6% qualified as AI high performers: organisations attributing at least 5% of Ebit to AI and describing the effect as significant. The figures measure different levels of impact, which helps explain how widespread employee benefit can coexist with a smaller reported effect on earnings.

That gap changes the conversation with boards because higher usage alone cannot establish sufficient return. “Organisations are saying, ‘I don’t know how to attribute value to AI’, and they’re also saying, ‘I’ve already blown past my budget for AI’,” Chitturu said. “That’s a very slippery slope for any technology.” Enterprises can therefore have productive individual users while still struggling to connect aggregate spending to earnings.

The difficulty of connecting spending to outcomes predates the latest generation of generative AI. Chitturu said CEOs in McKinsey client discussions had spent hundreds of millions of dollars on data and AI platforms since 2015 yet could not track what the spending had achieved. Those conversations have consequently shifted from selecting the next platform toward extracting value from investments already in place. The constraint is increasingly the relationship among cost, adoption and measurable business outcomes.

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.

The economic answer is to match the model and deployment mode to the workload

Once cost, adoption and outcomes are considered together, concentrated dependence on frontier-model suppliers becomes one architecture choice among several. Chitturu said organisations are asking whether they need to be “100% on these frontier AI companies like OpenAI and Anthropic” and exploring alternatives. “We are advising, and in fact their own CEOs are advising, that we need to diversify,” he said. Here, diversification means deciding what level of model capability each workload actually requires.

That decision starts with the task’s reasoning demands. Sophisticated adopters are classifying work by those demands, Chitturu said, sending high-order reasoning to frontier models and simpler, more deterministic tasks to open-weight models. Open-weight models make their weights available for an organisation to download, tune and run itself. Routine automation may gain little extra value from frontier reasoning, so an enterprise can assess whether the premium is justified for each class of request.

Chitturu puts the potential scope of routine work very high. “We need to create a solution that is fit for purpose because, honestly, 95% of enterprise activities don’t require a higher degree of computation than simple arithmetic. It is just an automation job.” The 95% figure is Chitturu’s assessment rather than an independently established share, but it illustrates the architecture he expects: greater computation is reserved for workloads whose reasoning demands justify it.

That workload split leads to an aggressive forecast from Chitturu and McKinsey. “In our view, the future of the enterprise AI architecture is going to be maximum 10-15% on the frontier, and 80-85% is going to be on open weights,” he said. The percentages describe a proposed segmentation of enterprise workloads, treating frontier capability as a premium resource within a broader model portfolio. As advice from an AI consultancy, that forecast also reflects McKinsey’s commercial position in helping enterprises design and change such architectures.

The economics behind that forecast become more compelling as the capability gap narrows. Chitturu said the delay before open-weight models reach frontier quality on most enterprise workloads has fallen from years to several months. He also puts their token price at roughly one-third to one-tenth of frontier API pricing. Those rates describe token prices; completed-task costs can still differ because models can consume different quantities of tokens.

Model access Approximate cost per million output tokens
Open-weight model through an API $1-$6
Frontier API $10-$50

Lower API prices are only one part of the deployment calculation because self-hosting adds fixed infrastructure expenses and machine learning operations, or MLOps, costs for deploying, operating and maintaining models. Chitturu said those fixed costs pay off only after workloads pass a sufficient volume threshold. Many enterprises are still beginning to develop the engineering capabilities needed to operate that environment effectively. The economic choice consequently depends on deployment mode as well as model class.

That deployment choice gives a mature architecture more than two options. High-volume routine traffic can favor self-hosted open weights when unit economics dominate, while domain-specific workloads can favor the same deployment approach when a company needs to tune a model on proprietary data. Workloads with irregular demand, low sensitivity and no strict data-residency requirement can instead use a hosted open-weight API. Hosted access avoids carrying fixed infrastructure for capacity needed only intermittently.

Difficult or high-consequence workloads create the remaining case for frontier APIs because model quality can justify the premium when reasoning demands or failure costs are high. The same segmentation addresses a major economic objection to open weights: access to cheaper model classes does not require an enterprise to operate every model itself. An organisation can combine self-hosting at sufficient scale, hosted open-weight services for suitable variable workloads and frontier services where their capabilities have the highest value.

The harness becomes an economic control plane

A mixed model portfolio creates value only when software can choose among those options while work runs. The relevant mechanism is the harness, the software surrounding a large language model, or LLM, or an agent that determines what context the model receives, how many tokens it can spend and how it approaches the task. Harness engineering has commonly focused on guardrails and risk, but Chitturu sees the same layer becoming a way to govern spending. The portfolio decision then becomes executable policy.

That policy turns model choice into a runtime architecture decision. “The harness gives you the flexibility to choose whether we go to that costly, token-eating LLM or do we go through an open-weight [model], and how do we dynamically adjust this on the fly,” Chitturu said. A request can be routed according to its needs while the harness controls how much context and token capacity it receives. Procurement defines the available choices; the harness decides how those choices are used for individual tasks.

Runtime choice creates an engineering problem because routing has to preserve enough capability for the task while controlling consumption. A harness can send routine requests to a cheaper model when premium reasoning adds little value, then escalate difficult or high-consequence cases to a frontier model. Token allowances provide a second cost lever by governing how much computation a task may consume. Together, model routing and token limits determine part of the workload’s cost before the task completes.

Those controls make harness engineering part of AI financial architecture as well as AI safety. The organisation must build the surrounding software and operational competence needed to make routing and token decisions reliably, which Chitturu says many enterprises are only beginning to do. Without that capability, a model portfolio with lower-priced options can still be difficult to operate economically.

ROI depends on workflow redesign

Once routing controls model cost, the ROI question shifts to how AI changes the work itself. McKinsey’s high-performer data points to workflow redesign as a major difference between high performers and other organisations. The same comparison shows differences in ambitions and measurement practices, linking process change to how organisations set goals and track results.

Practice AI high performers Other organisations
Fundamentally redesigned workflows 73% 25%
Had transformative AI ambitions 62% 17%
Tracked AI impact 40% 20%

The workflow gap matters because inserting AI into an existing process leaves many decisions about roles, handoffs and supervision unchanged. High performers more often redesign the process itself, allowing improvements in model capability to change who or what performs each step. Their higher rate of impact tracking then connects those process changes to outcomes. Individual productivity reports can complement that measurement, but they operate at a different level from organisational impact.

Chitturu describes the shift as organisational as well as technical. “They have all said that this is not only a technology transformation,” Chitturu said. “This is a transformation of the way we work. This is a transformation of the way we organise ourselves.” For leaders, the consequence is that architecture programmes and operating-model programmes increasingly have to move together.

That operating model also has to change as models improve. Chitturu said eight different LLMs set records on advanced-reasoning benchmarks during the last year, forcing organisations to reconsider repeatedly how much of each workflow a model can perform and how much human supervision remains useful. “This discussion is not happening as a one-time blueprinting exercise. That is happening every three months,” he said. Faster model progress turns workflow design into a recurring operating decision.

Human review shows what those recurring decisions look like in practice. McKinsey clients had previously discussed architectures in which AI checks AI while a person stays in the loop. Chitturu said newer client discussions ask whether that person is still required or whether human involvement can be reduced. Improvements in model capability can therefore change the allocation of work itself, with each review reconsidering the boundary between automated work and human supervision.

To manage that changing boundary, Chitturu proposes a dedicated “think tank” team. The team would monitor changes in solution architecture, bring in domain experts to decide case by case which work belongs with AI and which belongs with people, and follow the changing economics. That remit joins capability, workflow and cost because a task can move between people, open-weight models and frontier models as relative performance and economics change.

Executive teams need calibrated commitment

Continuous reassessment also gives executive teams a way to control how aggressively they deploy AI. Chitturu warned about companies deploying AI broadly across several business areas without a path for employees to adopt it. “Those organisations are going to fail,” he said. “They are going to burn money, and they are not going to realise the return on investment and lose conviction.” In his view, deployment breadth without adoption creates spending before the organisation has established a route to value.

Waiting carries a different exposure because competitors can keep moving while a cautious organisation holds back investment. Organisations that wait because they doubt the durability of AI advances can face what Chitturu called “high-growth, fast-moving disruptors who will come and completely decimate your industry”. He identifies the central challenge for executive teams and boards as deciding how aggressively to move and how much to commit. The decision is therefore about the pace and scale of commitment under changing technical and economic conditions.

That commitment has to remain adjustable because frontier APIs, hosted open-weight APIs, self-hosted models and human work carry different costs and operating requirements. Their relative value changes with workload volume, sensitivity, reasoning difficulty and the consequences of failure. Executive teams consequently need the ability to keep reallocating work and investment as models, economics and workflows change.

Key executive takeaways

  • Manage AI costs at the workload level: Falling token prices can still produce higher bills as agentic systems consume five to 30 times more tokens per task. FinOps teams need to track completed-workload costs alongside token prices and usage.
  • Connect AI spending to measurable returns: AI budget overruns and weak attribution are already constraining adoption, even as employees report productivity gains. Executive teams need metrics that connect AI consumption with workflow outcomes and financial impact.
  • Match models to workload requirements: Routine tasks can often run on lower-cost open-weight models, while complex or high-consequence reasoning can justify frontier APIs. Architecture teams can also choose between self-hosting and hosted APIs based on volume, sensitivity and operational capability.
  • Turn the AI harness into a cost control layer: Runtime routing and token limits allow enterprises to direct each request toward an appropriate model and control its compute budget. Platform teams need the engineering capability to apply these policies reliably as prices and model performance change.
  • Redesign workflows as model capabilities improve: McKinsey’s AI high performers report substantially higher rates of workflow redesign and impact tracking. Business and technology teams need recurring reviews of task allocation, automation and human supervision as models become more capable.
  • Calibrate AI investment with adoption and economics: Broad deployment without an adoption path can consume budgets before value materialises, while excessive caution can leave organisations behind faster-moving competitors. Executive teams need an adjustable investment model that reallocates workloads and spending as adoption, capabilities and economics change.

Alexander Procter

October 1, 2026

12 Min

Okoone experts
LET'S TALK!

A project in mind?
Schedule a 30-minute meeting with us.

Senior experts helping you move faster across product, engineering, cloud & AI.

Please enter a valid business email address.