The Agent Cost Reckoning Late last year, one of Uber’s roughly 5,000 engineers sat down for a coding session. Two hours later, the token meter read $1,200. By April, the rideshare giant had blown through its entire $3.4 billion AI budget for 2026. Not on training runs or exotic infrastructure—mostly on the quiet, daily consumption of coding agents that had become as routine as checking Slack. The math is unsettling and, by now, familiar inside enterprise IT. Token prices for GPT-3.5-level capability fell 98% in two years. Yet something perverse happened on the way to the bottom line. Enterprise AI bills swelled by an estimated 320% over the same window. Welcome to the era of agentic AI economics, where cheaper tokens meet insatiable consumption, and where your CFO has probably already stopped asking if an agent is capable, and started asking whether it earns its keep. A new McKinsey report, Is that AI agent worth it? Agentic economics and the modern operating model, puts a number on the anxiety: 93% of surveyed companies are exceeding their AI budgets, and one in five has already throttled usage just to keep operating costs under control.
McKinsey’s team flags a clean break in how we need to think about cost. “Per-token pricing has stopped being a useful measure for what enterprises actually pay for gen AI,” the report states. That’s not consultancy flourish. It’s a structural shift. When an agent loops repeatedly, summons tools, validates its own output, and re-injects entire conversation histories because models are stateless, the economic unit stops being the token and becomes the task completion. And the cost of a successful task swings wildly. Case in point: those agentic workflows consume roughly 1,000 times more tokens than a straightforward code-completion prompt or a chat turn. The appetite doesn’t stop. Goldman Sachs projects token consumption will climb 24-fold between 2026 and 2030, propelled largely by always-on enterprise agents. For a sense of the trajectory, AT&T’s internal AI systems now burn through 27 billion tokens a day—up from 1 billion just eighteen months ago.
Six Hands in the Wallet
McKinsey breaks the spending surge into six drivers, and while none of them are shocking in isolation, together they explain why the gap between the rate card and the budget is getting wider every quarter. First, there’s context bloat. Because LLMs don’t remember anything between calls, agents re-send accumulated state at every step. A ten-step agent run with a 30,000-token prompt can silently swallow half a million tokens. Developers on Hacker News have started calling it “the context tax,” and the April 2026 thread about Claude Opus 4.7’s new tokenizer—where token counts jumped up to 35% for identical text—racked up 621 points and 438 comments of existential dread. The second driver is the most quietly devastating: roughly 60% of an agentic task’s operating cost goes into response refinement, not the initial generation. Verification, correction, iteration—that’s where the meter spins fastest. The agent produces something, checks itself, fixes mistakes, and tries again. In multi-agent systems, some practitioners report that 30% to 60% of tokens are burned in loops that produce no outcome at all—just a digital conference room where the AI is essentially talking to itself. Then there’s the over-provisioning of reasoning. Enterprises routinely throw top-shelf models at tasks a smaller one could handle. It’s the Formula 1 car for a grocery run dynamic. As one Reddit builder on r/AI_Agents put it after Anthropic’s $50-per-million-output-token Fable 5 launch, “intelligent model routing isn’t an optimization anymore—it’s a production necessity.” Orchestration inefficiency, tool coordination cost, and information structure—particularly the way non-English text fragments into more tokens—round out the list. The language issue is a stealth budget incinerator. A Spanish prompt can consume 55% more tokens than its English equivalent, Arabic 230% more, and Japanese nearly 300% more. For multinationals, that turns a monthly $5,000 bill into $16,000 for no extra functionality. As one developer noted in a tiktoken benchmark thread, “this isn’t a model problem, it’s an infrastructure decision that shows up on the invoice.”
The Receipts Are Wild
Uber’s spring budget implosion isn’t a freak event—it’s a pattern. The company had roughly 5,000 engineers on AI coding tools, burning between $500 and $2,000 per person per month before management stepped in. A two-hour session hitting $1,200 was, apparently, not unusual. Uber’s response: a hard $1,500 monthly cap per employee per coding tool, with a real-time internal dashboard so engineers can see the cost burn rate themselves. The caps don’t pool across tools, so burning through one doesn’t affect another—an admission that the tools are essential, but somebody needed to put a governor on the engine. Microsoft pulled most of its internal Claude Code licenses six months after enabling them, having watched individual engineer bills climb into the same $500–$2,000 range. Around the same time, a single company allegedly racked up a $500 million Claude bill in one month after forgetting to set usage limits. That’s the stuff of nightmares and FinOps conference keynotes. J.R. Storment, Executive Director of the FinOps Foundation, captured the zeitgeist bluntly: “In April and May, I started hearing from companies: ‘Oh my god, we are 3x over our budget.’” Even the most bracing analyst numbers carry a surreal edge. Gartner now predicts that 40% of AI agent projects will be cancelled by 2027 on cost overruns alone—not capability failures, not security scares, just economics. Meanwhile, Forrester’s 2026 enterprise survey pegs negative ROI on 22% of agent deployments, not because the agents don’t work, but because infrastructure costs outpaced productivity gains. Premature deployment before infrastructure maturity could, they estimate, generate an average additional cost of roughly $2.1 million per organization.
When Agents Cost More Than People
It’s not just the licensing bill. AI agents are now taking actions in production systems that are irreversible, and the losses from bad decisions are starting to get priced. By 2025–2026, there were already documented cases of agents corrupting business-critical data with a single tool call. The cost here is asymmetric—a small error can cascade downstream in seconds. When an agent opts not to escalate to a human because it’s miscalibrated about its own confidence, those errors propagate silently. This is the “square cost curse” some researchers are starting to talk about: the price of a mistake isn’t linear; it compounds with the number of automated steps that follow. Trust remains fragile, and that directly curbs the ability to scale. Forrester’s consumer data shows just 15% of US adults trust companies that use AI for customer interactions. Three-quarters of UK online adults want to know when they’re speaking to a generative system. In B2B, the skepticism hits revenue directly: 19% of enterprise buyers said AI-generated inaccuracies or misleading results eroded their confidence in purchasing decisions. That’s not an abstract trust metric; that’s pipeline shrinkage. The reaction in some sectors has been to yank the AI back. A B2B SaaS company spent $280,000 on a “fully intelligent AI support system” in late 2025, watched customer satisfaction dive, and quietly reverted to a human-in-the-loop model. A freelance copywriter in the US got a rush order from a content agency to redo every page of a hotel website’s AI-written copy—20 hours, $2,000—which was exactly the budget the client had hoped to save by using AI in the first place. So-called “boomerang” hires are cropping up: employees laid off in the first wave of AI replacement are being called back to clean up the mess.
[SPONSORED]
COMFYUI WORKFLOW OPTIMIZATION
Reduce render times by 40% with our automated edge-silicon pipelines. Download Whitepaper.
Where the Money Should Go Instead
If the diagnosis is grim, the treatment is already taking shape in teams that treat cost engineering as a design discipline, not a cleanup step. Model routing is the nearest thing to a quick win. Cursor’s Router, trained on 600,000 real requests and evaluated on millions more in online A/B tests, saves around 60% of cost while maintaining frontier-quality output, depending on the workload. In its early-access phase, enterprises cut costs by 30–50% and saw no performance drop. Shanghai AI Lab’s Avengers-Pro takes a multi-model scheduling approach that achieves Pareto-optimal cost-performance frontiers—delivering GPT-5-medium-equivalent accuracy at 27% lower cost. The architecture isn’t magical; it’s just matching tasks to the cheapest model that clears the bar. As Cognition CEO Scott Wu put it on a forum, “you can use a model that still gets the job done—the uplift can be 5x or 10x” if you stop defaulting to the flagship model for everything. Treat context as a managed asset, not an afterthought. Teams that aggressively prune and cache context are reporting 50% to 80% drops in spend with no visible quality loss. The best operators separate identity (always-on rules about what the agent is) from behavior (on-demand instructions triggered by keywords). One engineering team that implemented layer caching and strict context budgets saw total token consumption drop 94% while cutting human agent work time by 87%. Change the unit of economic measurement. Several pricing experiments are moving from tokens or seats to actual business outcomes. Sierra, founded by OpenAI board chair Bret Taylor, charges only when an AI agent resolves a customer inquiry autonomously—transfers to a human are free. Salesforce’s Agentforce Help Agent now comes with a “pay per resolution” model. The logic is simple: when you pay for tokens, the meter only ever spins faster; when you pay for resolved tasks, the incentive flips toward efficiency. Don’t overlook the cost of human cleanup. The fastest way to make an agent project look cheap on paper is to ignore the downstream expense of correcting its errors. Organizations that measure total case resolution cost—AI interaction plus human escalation and rework—often discover the “cheaper” model was actually the most expensive. McKinsey’s own client work with a top pharmaceutical company showed that targeted Copilot deployment for sales insight and pre-call planning could yield a 1–2% revenue lift and cut content costs by up to 20%—but only when the organization measured the right outcome metric, not the per-seat cost of the tool. Smaller firms, different playbook. For mid-market companies, the path isn’t about building custom routers or negotiating volume GPU contracts. It’s about pick-one-workflow discipline: choose a single high-repetition process like sales quoting, order tracking, or supplier reconciliation, and solve it with a lightweight SaaS agent rather than aiming for “enterprise-wide intelligence.” Teams that start small, validate ROI on a single workflow, and only then expand, report that the equivalent of one or two employees’ monthly salary can turn into sustainable process improvement. Alibaba Cloud’s “Wanxiaozhi” and ShareQA’s multi-agent architecture are specifically aimed at lowering this entry bar, sidestepping the cloud-cost trap that kills pilots.
The Language Tax Is a Strategy Problem
It’s worth pausing on the tokenization gap, because it’s arguably the most invisible of the cost drivers. When Common Crawl is 46% English and BPE tokenizers compress English natively, the economic consequence is that the same product in Spanish, Arabic, or Japanese simply costs more to run. A single word like “implementación” fragments into four tokens. Over a million API calls, that inflates costs by thousands of dollars—pure linguistic overhead. The implication for global enterprises is stark: expanding AI services into new language markets carries a per-customer cost that grows with linguistic distance from English. It’s not just a localization challenge; it’s a margin problem that scales with user growth. Communities on GitHub are already building reproducible benchmarks to quantify the damage, and developers in Latin America and Africa are loudly noting that tokenization bias is a structural tax on digitization.
What Survival Looks Like
McKinsey’s recommendation boils down to a sober shift: the organizations that thrive won’t be the ones with the best agents, but the ones with the best agentic economics. That means building frameworks that track price per token, tokens per attempt, cost per successful task, and total organizational spend as four independent curves that can move in opposite directions. A more expensive model that gets it right in one shot can be cheaper per resolved task than a cheap model that triggers a cascade of retries and a human escalation. On the ground, the tools are emerging. Platforms like Speakeasy AI Cost Control give finance teams a unified view across Claude Code, Cursor, Codex, and others—with aggregation by employee, team, and tool. Lenovo’s Token Plan and AsiaInfo’s “Tianshu” token operations platform turn AI consumption into something approaching utility metering: real-time visibility, departmental chargebacks, hard limits. Microsoft’s Azure AI cost management dashboards now include budget alerts and anomaly detection specifically for agent workloads. Uber’s internal dashboard, crude as it may be, is a prototype of what’s coming. But tools alone won’t fix the underlying culture. As long as engineering teams treat frontier models as the default setting, budgets will leak. A mental model has to change: cost engineering needs to sit in the design review, not the monthly financial review. Reddit’s r/AI_Agents community has coalesced around a blunt mantra: “Cost engineering is now part of agent design.” It’s become table stakes.
The Big Question
None of this means agents are a bust. The technology works. The cases where agentic systems compress case resolution from hours to minutes are real. The numbers just don’t work if adoption keeps running on the assumption that tokens will trend to zero and consumption will be free. Jevons paradox is playing out in near real time: efficiency gains in model pricing are driving consumption so high that total spend explodes. The real wildcard is whether the 40% cancellation rate Gartner projects will be the result of genuine cost discipline—a healthy pruning—or whether it will be a panic reaction that kills good projects alongside the sloppy ones. For enterprises, the line between the two might be measured by one thing: whether they learn to budget by outcome before their CFOs force them to budget by cutoff date.