Red car and tire imprints on a beach as seen from above, Gruissan, France

The hidden cost curve in AI-assisted development

AI costs are spiraling — but the solution is improving cost visibility.


In brief
  • AI-assisted development costs are rising as token consumption grows with longer sessions, larger contexts and agent workflows.
  • Organizations need visibility into AI usage to identify inefficient workflows and control spending.
  • Seven key patterns can help organizations understand AI usage and identify where token consumption may be inefficient.

Software engineering firms are struggling with runaway AI costs, and for many companies, there appears to be no relief in sight. The industry’s situation today is a challenge for leaders because AI was predicted by many to be a cost-saver, not a cost-escalator. The thinking was that AI would transform the very structure of these companies, giving basic code-writing tasks to AI agents overseen by a small group of highly skilled human engineers. But the reality to date has been different.

 

AI-assisted development changes the cost profile for software engineering because costs are tied to usage shape, not tool access. As AI adoption has grown, so, too, have costs. Token consumption is exploding, usage patterns are inefficient, and consumption-based pricing makes spend unpredictable.

 

Today, some development teams are spending the equivalent of a junior engineer’s salary every month on tokens.1 And costs continue to spiral upward, to a level where AI spend is becoming a major operational issue at many companies.

 

It’s true that software development has never been inexpensive. Traditional software development tools have visible license costs. Build services have infrastructure costs. Cloud environments have metered resource costs. But AI development adds a new layer — inference costs created by prompts, context, tool calls, model routes, generated output, retries and agent behavior — all using tokens.

 

The hidden cost curve appears when normal developer behavior becomes repeated model consumption — ordinary workflow patterns that were never designed as economic decisions. As developers rely more on AI, sessions grow longer. Context grows. The assistant reads more files, tool calls multiply and the agent retrieves or verifies its own work. The final answer is reviewed, corrected and regenerated. None of these steps are unusual or unnecessary. In fact, they are normal components of AI development. That’s where the problem lies, because inference costs are difficult to both predict and manage, especially at companies where AI use is considered mandatory.

A simple cost mental model

Software engineering leaders should view AI-assisted development cost as the sum of several moving parts. In short, AI-assisted development equals session context:

  • Retrieved code and documents
  • Tool instructions and tool outputs
  • Model route cost
  • Generated output
  • Retries and repair loops
  • Verification steps
  • Cache writes minus cache reuse
  • Parallel agent activity

This mental model matters because it prevents leaders from focusing only on the listed model price. A cheaper model can still be expensive if the workflow sends too much context. A cache discount can still disappoint if cache reuse is low. A frontier model can be justified for hard reasoning but wasteful for routine file finding. A local model can reduce provider spend but still create quality risk if used for the wrong step.

When considering how to manage token costs, the economic question shouldn’t be “Which model is cheapest?” Leaders should be asking “Which segments of this development workflow are consuming tokens without producing enough trusted engineering value?”

For a CFO or FinOps owner, each component in this model is a cost line that can be measured and attributed to a particular team, project and repository. That is what turns AI development spend from a single provider bill into a managed operating line with visible unit economics.

Start with cost visibility

Before any single pattern can be controlled, the organization must see where tokens are spent. Teams cannot control what they cannot see. Many organizations can review total AI spend by provider. Fewer can see AI development consumption by product, repository, session, task type, model route, cache outcome, tool call pattern, retry count and quality result. Without that detail, leaders cannot tell whether spend is coming from productive development work, misrouted routine tasks, long sessions, agent loops or a small number of power users. The best approach is to consider instrumentation first. Measure the workflow before changing the platform.

Here are seven key patterns that software leaders should look for to better understand how developers are using AI — along with signals to watch to identify areas of improvement.

Pattern 1: Session rehydration

Long sessions are convenient for developers. They are expensive for AI systems. In many chat and agent workflows, later turns carry prior conversation, intermediate reasoning artifacts, tool outputs, file snippets and instructions forward. The longer the session runs, the more each new question can inherit old context. A small follow-up can become a large request because the system has to preserve continuity. This creates a simple pattern, where small requests, sometimes unrelated, use up more input tokens, and there is no session boundary, because debugging, planning, review and side questions blend into one cost pool. Long sessions aren’t always unnecessary. Some work needs repeated interactions with the agent and continuity is critical. But often there is no decision point that asks whether the next request should continue the session, compact the session, summarize prior work or start fresh.

Signals to watch:

SignalWhat it indicates
Rising average input tokens per turnSession history is becoming a cost driver.
High token cost for short user promptsOld context is dominating the request.
Many unrelated intents in one sessionThe assistant is carrying out unnecessary history.
Low cache reuse across similar sessionsPrompt or context structure is unstable.

Pattern 2: Context over-collection

Coding assistants need context. Uncontrolled assistants often collect too much of it. A developer asks for help with a bug. The assistant reads the ticket, searches the repository, opens several files, reads tests, inspects logs, asks for more context and then reads more files after the first attempt fails. Some of that work may be necessary. Some of it may be repeated because the assistant lacks a compact map of the codebase. The cost curve comes from the assistant reading broad code context when targeted code facts would be enough. This is where abstract syntax tree (AST) knowledge bases, repository maps, language-server metadata, dependency graphs and symbol summaries become economic controls. They let the system answer a smaller question before sending a larger context pack.

Signals to watch:

SignalWhat it indicates
High file read count per taskThe assistant may be exploring inefficiently.
Repeated reads of the same filesThe session lacks durable working memory or compact facts.
Large context with low edit surfaceThe assistant sends more codes than the change requires.
High retrieval volume with low task progressRetrieval is becoming activity rather than volume.

Pattern 3: Default frontier routing

Frontier models are valuable. They are also expensive. The cost issue appears when they become the default route for every step. Because model tiers differ in capability, cost and where they run, engineers should carefully consider where they request assistance. A local model running on the developer’s own machine is obviously going to be less expensive than a hosted cloud model. The key is identifying the class of work needed — file finding, summarization, draft planning, hard debugging, architecture reasoning, etc. and using the most appropriate model. For example, a workflow may need frontier reasoning for planning and debugging, but not for every search, classification, summary or formatting step. This is why route-by-step matters. The control question should be, “What is the next step?” not, “What is the default model for this assistant?” 

Signals to watch:

SignalWhat it indicates
High frontier route percentageExpensive models may be overused.
Frontier calls for low-complexity tasksRouting policy is too coarse.
Low retry rate on cheaper routesLower cost models may be acceptable for more task classes.
Quality failures after routing downA task class may need stronger model support.

Pattern 4: Tool and schema overhead

In this context, a tool is a specific capability the assistant or agent is allowed to call, such as reading a file, searching the repository, running tests, querying an application programming interface (API) or editing code. For each tool, the model is given a definition that states its name, purpose, inputs and expected output. Agentic systems often carry these tool definitions, schemas, instructions and examples so the model knows what it can do. That overhead can be useful. It can also become a hidden tax. If every request carries a large set of tool definitions, the model pays token cost before it sees the task-specific content. If tool outputs are verbose, each tool call can add more context to the next turn. If failed tool calls are repeated, the workflow pays for both the failed call and the extra model reasoning around it. The control goal is not to remove tools. It is to load the right tools at the right time and keep tool outputs proportional to the task. 

Signals to watch:

SignalWhat it indicates
High tool definition token shareTool setup is consuming budget before work begins.
Many tools loaded but few usedTool loading is too broad.
Large tool outputs followed by model retriesTool output is adding cost without resolving the task.
Repeated failed tool callsThe agent needs a stop rule or alternate path.

Pattern 5: Agent loops and parallel spend

Agents can work through tasks more independently than simple chat. That independence creates a new cost shape. An agent may plan, search, read, edit, test, inspect failures, revise and test again. That is useful when the agent is making progress. It is wasteful when the agent repeats the same reads, cycles through similar plans, retries failing tools or spawns parallel work without a clear budget. The issue is not a single expensive request. The issue is token velocity over time. That’s why agentic development needs progress checks. A useful control asks more than just if the request is allowed. It also asks, “Is this workflow still making progress worth the cost?” 

Signals to watch:

SignalWhat it indicates
High tokens per minuteToken velocity may be rising beyond expected task class.
Repeated intent across turnsThe agent may be looping.
High retry-to-success ratioThe system is paying for repair rather than delivery.
Many concurrent agents per user or repoParallel spend may need a budget guardrail.

Pattern 6: Output expansion

Output tokens can become expensive when assistants explain more than the task requires. Developers often need concise answers: the patch; the command; the failing assertion; the reason for a test failure or the next step. But AI tools often generate broader explanations or alternatives. Those outputs may be useful during planning but wasteful during routine execution. The problem becomes larger when verbose outputs are carried into later turns as session context. Output expansion becomes input expansion. For engineers, the solution is to ask AI to match answer size to task need. 

Signals to watch: 

SignalWhat it indicates
High output tokens for low-complexity tasksResponse policy is too loose.
Repeated clarifying turns after long answersVerbosity is not producing clarity.
Long outputs copied into later contextOutput cost is compounded into input cost.

Pattern 7: Cache without discipline

Caching is often treated as an automatic savings mechanism. It is not. Cache value depends on reuse. A prompt cache, context cache, exact cache or semantic cache can reduce costs when stable content repeats and when reuse is safe. It can also create disappointing economics when prompts are unstable, dynamic content breaks reuse or cache writes are not recovered through later hits. Cache should be measured as an economic system: write cost, read savings, hit rate, quality feedback and avoided provider calls. The control question should be “Is this cache paying back?” 

Signals to watch:

SignalWhat it indicates
Low cache hit rateCache design may not be worth the writing cost.
High prompt variationStable prefix design is weak.
Cache hit followed by correctionReuse quality may be poor.
No record of avoided callsSavings cannot be defended.

The compounding effect

These patterns rarely appear alone, which accelerates the cost spiral. As each decision affects the next, the cost curve compounds. A long session can carry too much context. That context can include broad file reads. The assistant can use a frontier model by default. The agent can load broad tool schemas, retry failed steps, produce long explanations, miss cache reuse and keep running because no progress check exists. Long sessions lead to larger inputs, higher model costs, larger outputs, more context in each turn, more expensive retries, weaker cache reuse and harder cost attribution. 

This is why AI development cost should be managed as a system. The enterprise does not necessarily need a cheaper model. It needs a way to shape the workflow before consumption happens.


Summary 

Smart companies will train engineers to better utilize the tools at their disposal. And most importantly, they will build in controls that can monitor the signals for these seven patterns to improve how AI is used across their development teams, ensuring that costs are managed without sacrificing speed and productivity.

Michael Flynn, EY US Technology Consulting Leader, contributed to this article. 

About this article

Authors

Related articles

The reckoning over AI cost and value has begun

The EY US AI Pulse Survey shows that executives must weigh the costs of action in AI against the price of inaction, to drive more than just adoption.

Shifting AI: experimentation to trusted output

AI works, but at what cost? Smart companies are moving to govern consumption without sacrificing developer speed

How can you break out of the AI ROI trap?

Tech companies can create value from AI investments by focusing on end-to-end transformation, fit-for-purpose KPIs and effective governance strategies.