Rising AI use and falling token spend tell a CTO little about whether AI creates value. Uber exposed the problem sharply: in December 2025 it gave engineers Claude Code and created internal leaderboards ranking teams by token consumption, yet by April its entire 2026 AI coding budget was exhausted. Uber President and COO Andrew Macdonald said the company still had no demonstrated connection between heavy use and better products for riders and drivers.
That mismatch has acquired the name “tokenmaxxing”: rapidly increasing token consumption without a corresponding return on investment. The financial stakes are growing because Gartner expects AI-agent-software spending to approach $207 billion this year, up more than 139% from $86.4 billion in 2025. Adoption and consumption are inputs; enterprise outcomes are shipped features, fixed bugs, solved customer problems, and other useful work whose value can justify the expense.
Those outcomes are hard to compare with spend because token expense varies with the work. The same engineer using the same tool on the same day can generate radically different costs when one task involves autocomplete and another sends parallel agents through a major database migration. Enterprise governance consequently has to answer two questions at once: where spending can be made more efficient, and which workloads produce results worth buying in the first place.
AI value requires outcome measures
Uber’s experience shows why adoption metrics become dangerous when they carry an implied value judgment. A leaderboard based on consumption rewards an observable input, while the company ultimately needs engineering outcomes. Near-universal AI adoption could coexist with little improvement in delivery, while a high-token workload could be highly productive. Consumption alone cannot distinguish those cases.
Because consumption cannot establish value, governance has to include allocation and outcomes. Leaders increasingly need to decide where expensive models and agentic workflows deserve the budget while establishing evidence that the resulting work matters. Cost-optimization mechanisms are already sophisticated enough to make some allocation decisions automatically. Measuring business return remains harder.
Caps serve different governance purposes
That allocation problem has already changed how companies control consumption. Uber’s immediate control is concrete: spending is now capped at $1,500 per employee, per agentic coding tool, each month. Microsoft questioned Claude Code license costs and canceled licenses across its Experiences and Devices division, while Duolingo abandoned a plan to incorporate AI use into performance reviews after employees objected to being pushed to use the tools for their own sake.
A cap can also act as a signal instead of simple rationing. Everlaw, which provides AI for litigation and investigations and therefore has a commercial interest in effective enterprise AI use, sets per-person token limits, but CTO Max Christoff uses them as escalation signals. When an engineer reaches the threshold, a one-line email goes to the tools team, and the engineer usually receives twice the allowance that same day. The event creates visibility and an opportunity for review without assuming heavy consumption is undesirable.
Agiloft, an enterprise contract lifecycle management platform with its own commercial stake in AI-enabled enterprise software, went further and eliminated its caps. Noe Ramos, its VP of AI operations, explained the decision: “74% of people were never hitting the old caps anyway. Caps were a false ceiling that created friction for heavy users without addressing the actual cost drivers.” For Agiloft, the restriction affected the minority with the largest workloads while doing little about how the underlying system selected models.
These different uses make “cost governance” imprecise unless leaders specify what a control should accomplish. Uber uses a fixed monthly limit; Everlaw uses a threshold to trigger attention and then readily expands the allowance; Agiloft decided its limits interfered with some heavier users while leaving the underlying expense mechanism intact. Because high-token work can be legitimate, a company can approve larger budgets, let individuals choose models, or route requests automatically through infrastructure.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.
Obvious waste can be removed before ROI is fully known
Once governance distinguishes legitimate heavy use from avoidable expense, infrastructure becomes an immediate target. Ramos sees infrastructure choices as a source of waste. “Teams aren’t burning spend because they love waste. They’re burning it because the default infrastructure pushes them toward it,” he said. “Most enterprises still hand model selection to whoever is prompting, which means a frontier model is handling tasks that a cheap open-weight model could do just as well. That’s not a people problem. It’s a plumbing problem.”
Promova, a language-learning company, encountered that problem in the defaults employees received. Head of engineering Dmytro Palaniichuk found premium models already selected on team and enterprise plans, with high reasoning effort enabled by default. On Promova’s plan, Opus sessions also moved to a 1M-token context window, and employees generally left these settings unchanged. Within months, Palaniichuk saw premium models being used for simple work such as checking email.
Those defaults concentrated Promova’s spending: Opus using the 1M-token window represented roughly one-third of its monthly expenditure. Promova is working toward the following model mix, although Palaniichuk treats the allocation as a direction rather than a permanent configuration.
| Model | Target share |
|---|---|
| Opus | 40% |
| Sonnet | 50% |
| Haiku | 10% |
The target makes model selection part of cost allocation by fitting the model to the work. Changing the default still leaves employee habits in the loop. “What’s stopping us is mostly a behavioral problem. Even if you force the organization default to Sonnet, but you don’t create awareness of what to use when, people fall back to habit,” Palaniichuk said. He argues that employees need to understand which model fits a task and build the habit of checking that choice before opening a new LLM conversation.
Agiloft has moved more of the same decision into infrastructure. It made cheaper models the default, escalates to frontier models when required, routes at the infrastructure layer instead of leaving routing inside individual prompts, and uses caching. Ramos calls a company-wide LLM gateway, the service through which model requests pass so policies can be applied centrally, the fastest initial improvement for many enterprises: “The quickest lift for most enterprise teams is a company-wide LLM gateway that handles cheaper defaults and smart routing.”
That approach leads Ramos to a stronger architectural conclusion: “Don’t ration the tool. Fix the architecture underneath it. Scarcity governance is a patch. Intelligent routing is the fix.” Merge, Databricks, AWS Bedrock, and Microsoft’s Azure AI Foundry are competing infrastructure offerings with auto-routing approaches designed to match task complexity with an appropriate model at an affordable cost. These providers have a commercial interest in wider adoption of routing infrastructure, while the operating mechanism they expose shifts more cost-performance decisions away from individual users.
Databricks shows how detailed that mechanism can become. At its Data + AI Summit in June, the company introduced Smart Routing inside Unity Gateway, alongside hard spending caps and cost attribution spanning hosted models, coding agents, and custom agents. Smart Routing is in beta and supports both recommendation-only and automatic-routing modes, letting an organization choose whether the system advises a user or makes the routing decision itself. Databricks sells this governance infrastructure, so it benefits commercially when enterprises adopt centralized routing.
David Nasi, Databricks director of product management, says Smart Routing evaluates each request with deterministic signals plus model-based classification. Its inputs include prompt intent and length, referenced files, stack traces, scope of change, required reasoning depth, and execution complexity. Together, those signals describe a task’s likely demands before the platform commits to a model and its associated cost.
Model choice is only part of that allocation because the surrounding execution setup also affects results. A harness is the software and configuration that controls how a model is invoked and given tools or context, and Databricks can select or configure it during routing. “We‘ve found that the same model performs differently when leveraged with a different harness, so we‘ve built Smart Routing to include that flexibility,” Nasi said. The relevant optimization unit can therefore include the underlying model and the way it is run.
Because automated allocation creates another decision layer, administrators also need to inspect what that layer does. Databricks exposes the runtime decision so administrators can see why routing occurred. “The decision is fully transparent at runtime; we explicitly avoid making the router a black box,” Nasi said. When spending reaches a budget ceiling, an administrator can block additional requests with a hard limit or configure a fallback that first tries a cheaper model that still meets the applicable requirements.
Agents make this control harder because one requested task can trigger dozens of model calls without a person approving each one. A per-call decision can consequently be too narrow for the actual unit of work. Databricks handles these workloads by evaluating at execution boundaries, meaning points where a meaningful stage of the agent’s work begins or ends, so the routing decision corresponds more closely to the larger task.
Early usage at Databricks suggests this kind of optimization can remove clear over-allocation. Teams that previously sent everything through frontier models are moving boilerplate generation, simple bug fixes, and minor edits to cheaper models without measurable deterioration in resolution rates, Nasi said. These mismatches can be corrected before an enterprise has a complete model of business ROI, leaving the harder value question to a separate layer of governance.
ROI requires measuring the work itself
Everlaw shows why that value question survives after cost allocation improves. Its quantified engineering projects put token expense beside estimated implementation effort:
| Everlaw project | Token spend | Engineering effort |
|---|---|---|
| Core Java infrastructure | $3,500 | Reduced from 9.5 engineer-months to 2.5 |
| Unlaunched product | $27,000 so far, likely to reach $40,000 | Reduced from an estimated 90–100 engineer-months to 19 |
The first case gives Christoff a concrete return against which token spending can be judged. “These are real numbers; the ROI of $3500 to save seven months of engineering time is simply a no-brainer,” Christoff said. The unlaunched product applies the same reasoning to a larger expenditure: its token bill is substantial, but the estimated reduction in engineering effort changes what that expenditure means.
Everlaw has also seen spending fail to create usable work. The company spent thousands of dollars having agents port interface code from Dojo to React, then discarded the generated output because the two frameworks have different assumptions about how state and view relate. The team revised the process so agents first document how the legacy system behaves, after which implementation can proceed from that behavioral documentation.
Together, the Everlaw cases establish a more useful unit for governance: useful outcome relative to expenditure. A substantial token bill can accompany a major reduction in implementation time, while another bill of thousands of dollars can end with code being thrown away. Raw consumption cannot tell a manager which situation is occurring, and reducing cost per call cannot make discarded work useful.
SUSE GM of technology and product Rick Spencer makes the same point through a hypothetical cost-return relationship. As an executive at a software company evaluating AI-driven development productivity, Spencer has a commercial and operational stake in how effectively such tools can be used. “The first reaction should not be to ‘use less AI’. High token usage can mean a lot of different things. If someone is using $16,000 in tokens to save us $100,000, of course we want to encourage that.” His operating rule follows directly: “It starts with diagnoses, not enforcement.”
That diagnosis begins by identifying the kind of work consuming the tokens. SUSE organizes usage into three categories managers use for coaching: Daily Work, Autonomous Agents, and Curve Jumping, the last covering one-time strategic efforts. Categorizing the workload gives managers a reason to consider what a model is doing before evaluating its consumption. Spencer also points out that the best model for an autonomous agent may differ from the newest frontier model, so a valuable workload can still contain opportunities for cheaper execution.
The categories support SUSE’s broader objective. “The important shift is that we’re not trying to suppress usage; we‘re trying to make sure usage maps to impact,” Spencer said. SUSE currently has no automated routing proxy, leaving these model decisions with individual developers, although a proxy is on its roadmap. Its evidence of value is currently anecdotal rather than expressed through quantitative AI ROI metrics.
Those anecdotes show the kinds of outcomes an evaluation system eventually needs to capture. One SUSE project reduced hundreds of dependency CVEs, Common Vulnerabilities and Exposures, to zero. Since May, its agents have also categorized nearly 10,000 CVEs in the VEX database. Results like these give managers an outcome they can relate to AI expense.
Promova demonstrates why establishing that relationship can still be difficult: optimization frequently ships alongside unrelated work. Palaniichuk could identify the large share of monthly expenditure associated with a particular premium-model configuration, but he could not isolate a clean before-and-after savings figure because the routing and model changes occurred with other changes. Cost telemetry can reveal a target for intervention while simultaneous changes make the intervention’s individual effect harder to isolate.
Databricks is building instrumentation aimed at closing part of that gap. Unity Gateway includes unified tracing, LLM-as-a-judge evaluation frameworks, evaluation datasets, trace analytics, and automated feedback loops. These tools can help a team observe requests, compare generated results, and detect quality changes after altering routing. For Databricks, those capabilities are also part of the commercial governance platform it sells to enterprises.
Evaluation makes cost optimization more measurable, but business-domain accuracy remains the customer’s responsibility. Palaniichuk describes the maintenance problem: “you’re imposing deterministic checks on non-deterministic output, and models ship on roughly a quarterly cadence, an eval tuned to one model doesn’t cleanly transfer to the next.” The evaluation layer therefore changes with the systems it evaluates, turning governance into an ongoing engineering activity.
Passing an evaluation still leaves another engineering concern because generated code has a useful life after the test runs. Tests can pass while the implementation remains difficult for engineers to maintain. Christoff describes one recurring failure pattern: “The coding agent will propose twenty surface-level fixes instead of addressing an underlying pattern.” A technically accepted output can therefore impose future engineering costs that a basic correctness test misses.
Because maintainability and architectural quality require local context, Everlaw gives engineers a broad menu of models and a dollar budget after models pass security review. The engineer who reviews the generated code can then develop practical judgment about which model tends to produce work worth keeping. That method preserves local discretion where a router or automated test may lack enough context.
The boundary matters because automation solves a narrower problem than ROI. Routing can assign cheaper resources, defaults can remove accidental premium consumption, caps can enforce or surface budget boundaries, and conscious model choice can improve individual decisions. Establishing whether the requested work created business value requires outcome evidence of the kind Everlaw can express through engineering-time figures and other teams are still learning to isolate.
Durable governance combines infrastructure, behavior, telemetry, and judgment
Because cost allocation and value measurement operate at different layers, practitioners disagree most clearly about where an organization should begin. Ramos starts with centralized routing because defaults and architecture can cause unnecessary expense before a user makes a meaningful choice. Spencer begins with coaching developers and training managers to diagnose types of usage. Christoff starts further upstream by asking what the business is optimizing for before selecting an intervention, while Palaniichuk begins by observing what teams actually do and identifying specific opportunities.
Palaniichuk’s discovery process gives teams tools and visible budgets to experiment across different models and providers. The organization then captures actual usage over a short period through telemetry and creates a feedback loop around the results. Teams use their own consumption data to find an appropriate allocation as they work, and Palaniichuk also warns that technical sophistication does not necessarily predict spending discipline.
Those observations explain why infrastructure and behavior can require separate controls. Ramos locates major avoidable cost in the underlying plumbing, while Palaniichuk finds that habits persist after defaults change. SUSE currently relies on coached developer discretion, and Everlaw gives engineers model choice within visible dollar budgets. Because these approaches govern different decision points, an enterprise can apply more than one as it learns where its own spending and value originate.
Promova’s model target illustrates why the governing discipline has to survive changes in individual models. “I don’t think this specific target will matter much, but the discipline will,” Palaniichuk said. Cheaper models will keep changing the available economics, while a growing choice of models increases the need to decide which tool fits a particular task. Christoff similarly expects today’s specific optimization tactics to expire even as the discipline behind them persists.
That changing environment makes cost optimization and value validation parallel activities. A company can fix an expensive default, route simple work to a cheaper model, or flag an unusual budget event immediately because those actions do not require complete ROI attribution. At the same time, telemetry, evaluation, code review, outcome measures, and business cases can accumulate evidence that the resulting work deserves the spending. Efficient allocation matters when the output remains worth producing.
Token budgets may become a business-planning category
Once token spending is tied to outcomes, the next change may be organizational. Christoff predicts that over the next one to two years token expense will increasingly be treated like annual headcount or production costs and less like general software or IT expenditure. Department leaders would enter annual planning with a position on both headcount and tokens, supported by a business case for each.
That planning model permits very different answers across teams. One department might rationally request millions of dollars in tokens while adding almost no employees; another could make the opposite choice. The resulting budget would express the department’s intended way of producing business outcomes, making token capacity part of the operating plan alongside the people expected to use it.
Main highlights
- Tie AI spending to outcomes: CTOs need measures such as engineering time saved, useful work shipped, bugs fixed, and customer problems solved to determine whether growing token consumption creates business value.
- Match caps to the governance goal: Fixed limits can contain spending, while escalation thresholds can surface unusual usage for review. Organizations need to decide whether a cap is intended to restrict consumption, trigger scrutiny, or support budget allocation.
- Remove infrastructure-driven waste: Platform teams can use cheaper defaults, centralized LLM gateways, caching, and smart routing to match model cost to task complexity. Transparent routing and fallback policies help preserve oversight as these decisions become automated.
- Measure the work AI produces: Engineering and business owners need to connect token expense with outcomes such as time saved, security issues resolved, and usable code delivered. Evaluation, telemetry, and code review can expose workloads that consume substantial resources without producing durable value.
- Combine infrastructure, behavior, telemetry, and judgment: AI governance needs to evolve as models, prices, and employee habits change. Central routing can address inefficient defaults, while training, visible budgets, evaluation, and local engineering judgment cover decisions infrastructure cannot make alone.
- Treat token capacity as an operating resource: Department heads can increasingly plan token budgets alongside headcount and production costs, with business cases tied to expected outcomes. This allows AI-intensive teams to justify higher consumption when the economics support it.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.


