AI coding token costs could soon rival developer salaries
Gartner expects AI token costs to meet or exceed the monthly salary of a typical software engineer within the next two years. The firm bases this forecast on a global average salary of about $2,000 per month. This does not mean AI costs will exceed every developer’s pay. In the US, software engineers can earn six-figure annual salaries.
The more important point is the change in the cost structure of software development. Enterprises are moving beyond fixed per-user software subscriptions. Generative AI and coding agents introduce a variable cost based on consumption. Every request, response and agent action can consume tokens, the units AI providers use to measure model input and output. Greater use can therefore create a larger bill.
The numbers can already become substantial. Nitish Tyagi, Senior Principal Analyst at Gartner, said he has heard cases such as, “My developer consumed $20K last month,” and “A business user consumed $32K.” These are anecdotal examples, not Gartner averages or forecasts. But they show the size of the exposure when consumption is not controlled.
This matters because AI agents can generate far more model activity than a developer making occasional requests to a coding assistant. An agent can execute multiple steps, collect context, generate code, review results and repeat work. Long context windows also increase the amount of information processed. As companies give agents more tasks and autonomy, token consumption can rise without an equivalent increase in business value.
For the C-suite, AI inference should therefore become an explicit part of software engineering economics. A fixed software license is relatively easy to budget. Consumption-based AI is not. Finance and engineering teams need to know the cost of AI by developer, workflow, product and business outcome.
The goal is not to restrict useful AI. It is to stop uncontrolled consumption. Tyagi said Gartner’s purpose is “to alarm the industry about the impact of token cost if it is not governed and controlled.” That distinction is important. A high AI bill can be rational when it produces sufficient value. High consumption without measurable value is the problem.
Executives should expect AI costs to become a material component of engineering budgets. This changes the investment question. Leaders should no longer ask only how much an AI coding product costs per seat. They also need to ask how much model consumption its workflows create, which tasks justify that expense and what measurable business result the additional spending delivers.
Enterprises lack the visibility and governance needed to control AI coding costs
The main constraint is not access to AI. It is cost control. Enterprises are scaling AI coding agents faster than their governance and financial systems are adapting to consumption-based computing.
Token costs are highly variable. A simple task may require limited model usage. An agentic workflow can involve repeated calls, growing context, retries and several stages of automated work. This makes the eventual cost harder to predict from the number of employees using the product.
Visibility remains another problem. According to Tyagi, companies often lack transparency into how token consumption is calculated and billed. AI coding vendors also have yet to provide what he calls “mature, built-in cost optimization capabilities.” Enterprises therefore cannot assume the software itself will prevent inefficient consumption.
The commercial model adds pressure. AI providers must continue investing in models and infrastructure while building profitable businesses. Gartner expects this dynamic to put further pressure on pricing. At the same time, adoption is expanding beyond developers. Business users who initially use AI lightly may consume more tokens as these tools become part of their daily work.
Internal management systems are also immature. Many companies cannot yet connect token spending to a reliable measure of return. Agent workflows can use more context than necessary. Automated processes can consume budgets earlier than planned. Teams can then struggle to explain whether the additional spending produced better software, faster releases or higher customer value.
This is why AI cost management must operate at the workflow level. A monthly total is useful for accounting but arrives too late for operational control. Engineering leaders need usage thresholds, automated monitoring and clear escalation rules. They also need to identify the users, models and workflows responsible for unusually high consumption.
Cost governance should not mean imposing the same token limit on every task. A high-value engineering problem may justify significant model spending. Routine work may not. The relevant measure is the value produced relative to the resources consumed.
Tyagi captures the executive risk clearly: “Without a governed engineering operating model, costs can escalate faster than the productivity gains these tools are designed to deliver.” Companies should respond by making token economics part of engineering management now. AI coding can create substantial value, but scale without visibility turns a technical productivity tool into an uncontrolled variable cost.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.
More token consumption does not mean higher developer productivity
There is no direct relationship between the number of AI tokens a developer consumes and the productivity gained, according to Nitish Tyagi, Senior Principal Analyst at Gartner. This is a critical distinction. More AI activity does not automatically create more business value.
Tyagi puts the issue plainly: “Tokenmaxxing is not directly related to higher productivity gains, but optimizing token consumption is.” An AI system can process large amounts of information, generate several responses and perform repeated agent actions without materially improving the final result. Companies that measure AI adoption through usage alone risk rewarding consumption rather than outcomes.
Context is one source of unnecessary cost. AI coding tools often need information about source code, documentation, requirements and previous interactions to complete a task. But larger context is not always better. Irrelevant files, duplicated information and lengthy histories increase the number of tokens processed. They can also make it harder for the model to focus on the information that matters.
Gartner therefore recommends context engineering. This means deliberately deciding what information an AI model receives and how that information is structured. Developers should include relevant material, summarize it where possible and remove unnecessary data. The aim is to use enough context to produce a high-quality result without paying to process information that adds little value.
The economic impact becomes more important as enterprises deploy autonomous agents. An agent may perform multiple model calls to plan a task, inspect code, generate changes and evaluate its own work. Poor context management can increase consumption at every stage. Small inefficiencies can therefore become significant when repeated across large engineering organizations.
For executives, the correct target is productivity per unit of AI spend, not token volume. That requires linking consumption to outcomes such as development time, software quality and feature delivery. A team that uses fewer tokens to produce the same result has improved efficiency. Higher spending is justified only when it creates enough additional value.
Gartner’s broader productivity estimate provides useful context. Tyagi said AI-assisted development can deliver productivity gains of up to 20%. He described this as “not a bad number.” The source does not establish that context engineering itself produces that 20% gain, nor does it claim every organization will achieve it. The figure instead shows that meaningful productivity improvements are possible when AI is deployed appropriately.
This makes context engineering both a cost discipline and an engineering skill. Tyagi advises developers: “Target context engineering as one of the most important skills for yourself. This is not only going to help your employer, but also your career.” For management teams, the implication is clear: training and operating standards should focus on efficient AI use, not simply greater AI use.
Lines of code is no longer a useful measure of developer productivity
AI can generate large quantities of code almost immediately. Gartner notes that it can produce entire Python libraries. This makes lines of code written an increasingly weak measure of engineering productivity.
The problem is fundamental. Code volume measures output quantity, not business value. More generated code does not establish that a product is more reliable, reaches customers faster or solves a more important problem. AI makes this weakness much more visible because software teams can now increase code output without a proportional increase in human effort.
Tyagi recommends measuring value through quality, speed and customer satisfaction instead. These metrics move the focus from how much code developers produce to what the organization achieves with it.
Release speed is one useful measure. Leaders can track how quickly teams move important features into production. They can also measure the time between application development and feedback from business, product and engineering teams. Shortening that cycle can help teams identify problems sooner and respond to changing requirements faster.
Quality must remain part of the equation. Faster code generation has limited value if it creates defects, security problems or additional maintenance work. Executive dashboards should therefore combine delivery speed with measures appropriate to the organization, such as production defects, reliability or rework. These additional metrics are sensible management measures, though they are not specific statistics cited by Gartner in the source article.
Customer outcomes matter as well. Tyagi argues that shipping features quickly while maintaining quality can strengthen competitive advantage and improve user and customer experience. This provides a more useful basis for evaluating AI investment than counting generated code or tracking raw token consumption.
The shift also changes how executives should assess developer performance. AI may allow an engineer to produce substantially more code, but that does not necessarily mean the engineer is proportionally more productive. The relevant questions are whether important work reaches users sooner, whether quality is maintained and whether collaboration and feedback become more efficient.
For C-suite leaders, AI creates an opportunity to replace activity-based engineering metrics with outcome-based ones. Token consumption and code volume remain useful operational signals, particularly for cost management. They should not be treated as measures of value. The executive objective should be better software delivered faster at an economically justified cost.
Formal governance is necessary to prevent uncontrolled AI spending
AI token costs are variable by design. That makes governance an operating requirement, not an optional finance exercise. Gartner recommends three basic controls: token thresholds, automated usage monitoring and explicit escalation policies.
Token thresholds give teams a defined spending boundary. Automated monitoring shows when consumption is approaching or exceeding that boundary. Escalation policies determine what happens next, including who reviews unusual activity and when additional spending requires approval. Gartner states that “embedding these controls into engineering workflows ensures consistency and prevents uncontrolled cost growth.”
These controls should operate close to the point of consumption. Aggregate monthly spending tells executives how much the company has already spent. It does not explain which workflows produced the cost or whether those workflows created enough value. More detailed monitoring can help engineering and finance teams identify high-consumption applications, agents and tasks before costs become difficult to reverse.
Governance also needs to address autonomy. Gartner recommends a “use case driven” framework that determines when a coding agent should be used and how independently it should operate. The firm proposes three execution models: developer-led, developer-with-agent and fully agent-led. The appropriate model depends on the task.
This distinction matters because greater autonomy can create more model activity. A fully agent-led workflow may plan work, collect context, execute tasks and perform repeated AI calls with limited human intervention. That can be useful for suitable workloads, but it also changes the organization‘s cost and control exposure. Enterprises should therefore make autonomy a deliberate engineering decision rather than a default setting.
Executives should avoid imposing a single consumption policy across every engineering task. A complex, high-value project may justify substantially more AI spending than a routine task. The goal is to define where additional autonomy and consumption produce sufficient value, then apply tighter constraints elsewhere.
Governance also cannot depend on developers making individual cost decisions. Nitish Tyagi, Senior Principal Analyst at Gartner, notes that developers tend to optimize for speed and convenience rather than cost efficiency. That behavior is rational when their primary responsibility is shipping software. Management systems must therefore make cost discipline part of the workflow rather than expecting every engineer to continuously calculate token economics.
The executive requirement is clear: assign ownership for AI consumption and connect technical controls with financial accountability. Engineering needs enough flexibility to capture AI’s productivity benefits. Finance needs enough visibility to forecast spending. Business leaders need evidence that higher consumption produces better outcomes. Governance should connect all three.
Use smaller AI models when the task does not require frontier-level capability
Model selection is one of the most practical controls for AI costs. Gartner recommends choosing models according to task complexity instead of sending every workload to the most advanced model available.
The reason is economic. Different engineering tasks require different levels of model capability. Routine and high-frequency work may be handled effectively by smaller models. Complex work may require a frontier model, meaning one of the most capable models available. Paying for frontier-level capability on every request can increase costs without creating a corresponding improvement in output.
Gartner recommends breaking work into smaller tasks that can be handled by smaller models, with “escalation only when complexity demands it.” Engineering teams should deliberately route simple, frequent work toward smaller models and reserve frontier models for complex, high-value workloads.
Task decomposition is important here. A large development request can contain several operations with different levels of difficulty. Some may require advanced reasoning or extensive context. Others may be simple and predictable. Separating those operations allows an enterprise to use expensive model capacity only where it contributes additional value.
This approach also provides executives with a better way to manage AI economics. Instead of asking which single model the organization should standardize on, leaders can establish a model portfolio and routing policy. The relevant questions become: What capability does this task require? What is the lowest-cost model that can meet the required quality level? When should the workflow escalate to a more capable model?
Cost cannot be the only selection criterion. Smaller models can reduce consumption expense, but using an inadequate model can create errors, retries or human rework. Those costs can eliminate the initial saving. Model selection should therefore account for total task economics, including model charges, execution time, output quality and the cost of correcting failures.
Security, data handling and reliability also remain enterprise requirements. A cheaper model is not the correct choice if it does not satisfy the organization‘s security, compliance or performance standards. These considerations extend beyond the specific Gartner recommendations in the source, but they are necessary when enterprises turn model routing into production policy.
Gartner’s recommendation ultimately supports a simple operating principle: match computing resources to the value and complexity of the work. Frontier models should be used where their capabilities materially improve results. Smaller models should handle tasks they can complete reliably. This gives enterprises a direct way to reduce AI costs while preserving the quality that justified adopting AI in the first place.
Context engineering and token reviews should become standard engineering practices
Context engineering is becoming a core skill for controlling AI quality and cost. It means deciding what information an AI system receives, keeping that information relevant, and removing content that does not help the model complete the task. Gartner recommends making these practices part of normal software development rather than leaving them to individual preference.
The financial impact is straightforward. AI models process the context developers and agents send to them. Long source files, repeated documentation, unnecessary conversation history and irrelevant data can increase token consumption. In agentic workflows, this overhead may occur repeatedly across multiple model calls. Better context management reduces that waste while preserving the information needed to produce a useful result.
Gartner recommends that developers include only relevant information, summarize content where possible and eliminate unnecessary data. This requires more than writing shorter prompts. Teams need to understand which code, requirements, documentation and previous outputs are necessary for a specific task. Context engineering should therefore be treated as an engineering discipline with defined practices and training.
Regular reviews are the second part of the operating model. Gartner advises embedding token-usage reviews into development cycles. Teams can identify workflows with unusually high consumption, investigate why they require so many tokens and determine whether the cost is producing enough value. This creates a repeatable process for refining AI usage as models, tools and workloads change.
The review should connect technical consumption with business outcomes. A high-token workflow is not necessarily inefficient. Complex work can justify higher spending when it delivers sufficient quality, speed or customer value. The useful question is whether the same outcome could be achieved with less context, fewer model calls, a smaller model or a different level of agent autonomy.
This cannot depend entirely on individual developer behavior. Nitish Tyagi, Senior Principal Analyst at Gartner, notes that developers tend to optimize for speed and convenience rather than cost efficiency. Enterprises should expect this because engineering teams are normally measured on delivery. If management also wants cost-efficient AI use, it needs to make token discipline part of engineering standards, tooling and review processes.
Tyagi also sees context engineering as an important professional capability. His advice to developers is explicit: “Target context engineering as one of the most important skills for yourself. This is not only going to help your employer, but also your career.”
For executives, the priority is to institutionalize this knowledge. Establish context standards, train developers, track token consumption and review expensive workflows regularly. The objective is not simply to cut tokens. It is to reduce consumption that does not improve the result.
Rising token costs are a reason to improve AI operations, not abandon AI coding
Higher AI spending does not invalidate the business case for AI coding tools. It changes the standard by which deployments should be managed. Gartner’s position is that enterprises should optimize AI consumption without sacrificing the value these tools can create.
Nitish Tyagi, Senior Principal Analyst at Gartner, cautions leaders against responding to rising costs by withdrawing from AI coding agents. He also advises against shifting every workload to open generative AI models purely as a cost response. His principle is clear: “The goal is always to optimize costs without compromising the value.”
Open models can be appropriate for some workloads, but their economics cannot be judged only by access or model costs. Enterprise deployment may also involve infrastructure, operations, security, monitoring and specialist skills. The correct choice depends on the workload and the total resources required to deliver an acceptable result. The source article does not provide a direct cost comparison between open and commercial models, so it does not support the conclusion that either approach is inherently cheaper.
Companies should instead start with their engineering maturity and the work they want AI to perform. Tyagi recommends starting small and focusing first on context engineering. Organizations should then choose the appropriate level of agent autonomy. Teams with immature controls should not assume that maximum autonomy will produce maximum value.
There is a measurable productivity opportunity. Tyagi says AI-assisted development can provide productivity gains of up to 20%, describing that result as “not a bad number.” The figure should not be read as a guaranteed return for every company or workload. It demonstrates the potential Gartner sees in assistive AI development and reinforces the need to compare productivity gains with the cost required to achieve them.
The executive decision is therefore not simply whether to adopt AI coding. It is where AI produces enough value to justify its variable cost. That requires measuring feature delivery, quality and other relevant business outcomes alongside token spending. It also requires different controls for different use cases.
Organizations can then expand successful deployments with evidence. Workflows that deliver meaningful gains at an acceptable cost can receive more investment or autonomy. Workflows that consume heavily without improving outcomes should be redesigned, moved to a different model or constrained.
Tyagi summarizes the core risk: “Without a governed engineering operating model, costs can escalate faster than the productivity gains these tools are designed to deliver.” The opportunity remains substantial, but the operating model matters. Enterprises that treat token consumption as a managed engineering resource will be better positioned to scale AI coding while protecting its economic value.
The bottom line
AI coding is changing the economics of software development. Gartner expects token costs to reach or exceed a typical developer’s monthly salary, based on a $2,000 global average, within two years. The executive response should not be to restrict AI. It should be to manage consumption with the same discipline applied to other material technology costs.
The priority is visibility. Leaders need to know which teams, agents and workflows consume tokens, what models they use and what business outcomes that spending produces. Token volume and lines of code are not measures of value. Faster delivery, software quality and customer outcomes are.
Cost control should then move into engineering workflows. Set consumption thresholds. Monitor usage automatically. Match models to task complexity. Define when agents can operate autonomously. Train developers in context engineering and regularly review high-consumption workflows.
AI-assisted development can deliver productivity gains of up to 20%, according to Gartner Senior Principal Analyst Nitish Tyagi. Capturing that value requires keeping the economics under control as adoption scales. Companies that establish this operating discipline early can expand AI use with evidence rather than assumptions and keep spending aligned with measurable results.
A project in mind?
Schedule a 30-minute meeting with us.
Senior experts helping you move faster across product, engineering, cloud & AI.


