Revamping Traditional IT Budgets: The Practical Fix for Exploding AI Spending

Illustrated armored robots counting cash at a chaotic office desk representing runaway enterprise AI costs and budget overruns

Summary

Enterprise AI costs are spiraling, and the root cause is structural: budgets built for predictable per-seat SaaS pricing were never designed for open-ended, per-token inference. Unmonitored AI agent loops, self-corrections, and chained calls generate token volumes no IT budget was built to anticipate, and the shift to agentic workflows can multiply that consumption a thousandfold. Erhan Giral of BMC Helix sets out why the two biggest culprits are unmonitored agent loops and context re-sends, and makes the case for intelligent model routing – a governance layer that sends routine work to smaller models and escalates to frontier models only when the task demands it – as the practical fix.

For many organizations, enterprise artificial intelligence (AI) spending is spiraling out of control. AI costs have emerged as the number one pricing and governance issue for enterprises running AI at scale this year.

AI Budget Issues 

The problem is two-fold: frontier model pricing is rising fast, and the cost of running AI models at enterprise scale is outsized for traditional technology budgets. Enterprises are learning painfully that buying AI and running it at scale carry very different costs. The root cause: AI budgets were never designed for the way AI actually operates. It’s a disconnect companies are scrambling to solve. 

AI inference is typically billed per token. Unmonitored AI agent loops and redundant context re-sends produce costs no budget was designed to anticipate. This is especially problematic for the IT industry, where budgets are built for predictable per-seat or per-license SaaS costs, not open-ended per-token consumption.

This problem is exacerbated by the shift toward agentic AI, where workflows can multiply consumption a thousandfold. The IT industry has seen this movie before: FinOps emerged when cloud’s consumption-based costs outgrew traditional budgeting. Now C-suite leaders are under pressure not only to prove that business outcomes warrant the investment, but to get costs under control quickly. 

Examples of AI Cost Management

Tesla has tried to get costs under control by capping AI spend at $200 per week, per employee. After Uber blew through its entire 2026 AI coding budget in four months, driven largely by Claude Code, the company responded by capping employee spend at $1,500 per month, per tool. Microsoft went so far as to cancel most of its internal Claude Code licenses, redirecting engineers to its own GitHub Copilot CLI (Command-Line Interface). And Meta is building real-time cost monitoring after projecting that internal AI spend will reach billions by year’s end. These stories are not outliers, but early signals of a systemic failure that must be addressed.

The Solution to AI Cost Management

Enterprises need a smart router: logic that sends queries to frontier models only when absolutely necessary. Strategies like this are part of an AgentOps framework – an emerging enterprise control layer that governs, observes, orchestrates, quantifies the value of, and optimizes digital labor at scale. Enterprises can start implementing it today. Here’s how:

  1. The first step is to name the structural flaws driving runaway costs: token pricing obscures the true cost per task, and CFOs lack the visibility to forecast or cap inference spend in real time. 
  2. The second is to close the governance gap around agentic AI, which comes down to two cost levers. 

The Two AI Cost Levers

The first lever is unmonitored AI agents: retries, self-corrections, and chained calls generate enormous token volumes that companies, leaders, and employees never see. That cost compounds when a single autonomous workflow can fire thousands of calls in minutes. 

The second lever, and perhaps the biggest culprit, is context re-sends. In May, Stanford’s Digital Economy Lab reported that input tokens, not output, are the primary cost driver in agentic workflows, and that agentic coding tasks consume roughly 1,000x more tokens than comparable chat interactions. 

Addressing AI Cost Issues

To fix both flaws, IT leaders should shift from cost-per-token metrics to cost per completed task, tying infrastructure metrics to operational outcomes. 

Intelligent model routing is the primary vehicle for that shift. A routing layer classifies each request by complexity and risk, sends routine work to smaller models, escalates to frontier models only when the task demands it, and falls back automatically when output quality checks fail. 

Most enterprise AI volume still defaults to the most expensive frontier models whether the task requires it or not, and peer-reviewed routing research has demonstrated inference cost reductions of up to 85 percent while retaining roughly 95 percent of frontier-model quality.

The Path Forward for IT Leaders

  • Audit current AI inference spend and identify where AI agent loops, context re-sends, and single-model defaults are inflating costs
  • Pilot intelligent routing on the highest-volume, lowest-complexity tasks first
  • Establish cost per completed task as the primary AI efficiency metric in executive reporting
  • Give service management and FinOps joint ownership of AI consumption, and hold AI agents to the same observability and anomaly-detection discipline AIOps already applies to infrastructure.

The promise of agentic AI to drive new business capacity is within reach. But it requires enterprises to rewrite their operational playbooks and govern agentic workflows for how they actually behave.

AI Costs FAQs

Why are enterprise AI costs exceeding IT budgets?

Traditional budgets are built for predictable per-seat or per-license SaaS costs, but AI inference is typically billed per token. That structure was never designed for open-ended consumption, and agentic AI workflows can multiply it a thousandfold.

What are the two biggest drivers of runaway AI costs?

The first is unmonitored AI agents: retries, self-corrections, and chained calls that generate enormous token volumes nobody tracks. The second is context re-sends. Stanford’s Digital Economy Lab found that input tokens, not output, are the primary cost driver in agentic workflows, and that agentic coding tasks consume roughly 1,000 times more tokens than comparable chat interactions.

How have companies tried to control AI costs so far?

Tesla capped AI spend at $200 per employee per week. Uber capped spend at $1,500 per employee per month after exceeding its 2026 AI coding budget in four months. Microsoft cancelled most of its internal Claude Code licenses in favor of GitHub Copilot CLI, and Meta is building real-time cost monitoring as its internal AI spend heads toward billions.

What is intelligent model routing?

It’s a routing layer that classifies each request by complexity and risk, sends routine work to smaller, cheaper models, escalates to frontier models only when the task demands it, and falls back automatically if output quality checks fail.

How much can intelligent routing save?

Peer-reviewed routing research cited in this article found inference cost reductions of up to 85 percent while retaining roughly 95 percent of frontier-model quality.

What should IT leaders do first to get AI costs under control?

Audit current AI inference spend to find where AI agent loops, context re-sends, and single-model defaults are inflating costs, then pilot intelligent routing on the highest-volume, lowest-complexity tasks first.

Erhan Giral
Erhan Giral
VP of AI Strategy and Innovation at BMC Helix

Erhan Giral is VP of AI Strategy and Innovation at Helix, with fifteen years in software intelligence and multiple patents in operational intelligence and application monitoring. He specializes in scalable AI architectures, knowledge graph-based reasoning, and machine learning platforms for observability and IT service assurance, with a particular focus on automated diagnostics and self-remediation for large-scale IT and telecom systems.

Want ITSM best practice and advice delivered directly to your inbox? Why not sign up for our newsletter? This way you won't miss any of the latest ITSM tips and tricks.

nl subscribe strip imgage

More Topics to Explore

Leave a Reply

Your email address will not be published. Required fields are marked *