In May 2026 McKinsey surveyed enterprise AI FinOps practices across five major industries — 120 enterprise participants, 75 of them qualified respondents — and the headline is stark even for a small sample: 93% report exceeding their AI budgets (McKinsey Enterprise AI FinOps survey, May 2026). Most were not over by a rounding error. Nearly half reported overruns of 10–30%, and organisations moving from isolated use cases to enterprise-wide adoption saw AI spend increase nearly fourfold.
The stranger number sits underneath the overrun. In McKinsey’s experience, 20–30% of enterprise AI spend is often unaccounted for — scattered across cloud providers, foundation-model contracts, AI-enabled software features, experimentation environments and business-unit purchases, with no single source of truth. A budget cannot govern spend the organisation cannot see. The overrun is not the disease; it is the symptom of a missing record.
Consumption pricing removed the safety net
Traditional software pricing hid usage. A seat cost the same whether it was touched once or a thousand times, so demand could sprawl without financial consequence — the bundled fee acted, in McKinsey’s phrase, as a safety net that hid the usage costs. Consumption pricing removes it. With model providers charging by usage, every prompt, retry, and agent loop lands on the bill.
Demand itself is also harder to predict than anything technology budgeting has dealt with before. The same task can generate dramatically different token volumes — research McKinsey cites puts the variation at up to 30 times (Stanford Digital Economy Lab, via McKinsey, Jul 2026). Agentic workflows multiply model calls per interaction. And employees who could never have built software now stand up AI applications and agents with little engineering effort, so consumption appears in places no budget line anticipated.
The capabilities that would catch this early are mostly not built. In the same survey, only 20–25% of companies report mature AI FinOps practices: 25% for spend visibility, 24% for token-consumption tracking, 21% for cost allocation, and 20% for forecasting and budgeting.
Govern the outcome, not the token
The instinct under a 93% overrun is to cut. McKinsey’s argument — and ours — is that cutting is the wrong first move, because the problem is not that the spend is too high. It is that nobody can say what the spend buys. McKinsey’s recommended unit of governance is the completed business outcome, not the token cost: cost per claim processed, cost per case resolved, revenue per AI-enabled workflow.
That reframing has a name worth noticing. McKinsey describes an emerging AI control plane — the management layer between users, applications, and the models they consume — and suggests that if ERP systems became the system of record for financial transactions, AI control planes may become the system of record for intelligence consumption (McKinsey Quarterly, Jul 2026). Whatever the label, the function is the point: one place where spend, usage, and business outcomes meet, so that a budget conversation can be about value rather than variance.
Discipline shows up in the survey’s own terms. Organisations with high forecasting maturity save 10% more on AI spend than their peers on average, and McKinsey’s experience is that companies thoughtful in their AI consumption can save 20–30% of their AI costs. One figure is from the survey and one from engagement experience; neither is a promise about any particular organisation.
What would count as proof?
- A dated, single view of AI spend that spans cloud providers, model vendors, software platforms, experimentation environments and business-unit purchases, with the sources it was assembled from listed.
- Attribution of consumption to the business unit, workflow, and named owner that generate it, at a granularity finance accepts.
- A demand forecast with its adoption, workload, and pricing assumptions stated, and a record of forecast against actual for at least one budget period.
- A defined cost-per-outcome metric for each material workflow — the completed business outcome named, the cost assembled from usage records, the calculation reproducible.
- For any claimed saving from optimisation, the specific action taken, the before-and-after usage records, and the invoice line that moved.
- An owner for the record itself: the team accountable for keeping spend, usage, and outcome data joined as models, prices, and workloads change.
The test mirrors the budget meeting nobody enjoys. When the number comes in high, can the organisation say which workflows spent it, what those workflows returned, and what it will choose differently next quarter? If the answer is a vendor invoice, there is no answer.
What remains unclaimed?
These are survey findings and practitioner estimates, not laws. A 93% overrun rate among 75 qualified respondents in five industries describes that sample, not every enterprise. The 20–30% of unaccounted spend is McKinsey’s experience across engagements, not a measured global figure. Savings attributed to optimisation are self-reported, and the survey’s own exhibit shows most savings realised in its three-month window came in at 20% or under.
Nothing here shows that visibility alone reduces spend, that a control plane pays for itself, or that any particular organisation’s unaccounted share is 20–30%. And a spend record, however complete, is half a case: it prices the demand without valuing the work. The other half is the value record — a different discipline, and the one the budget conversation is actually waiting for.
Count the spend where the value is counted
The overruns will keep coming as long as demand is unmetered and ownerless. The organisations that get ahead of them will not be the ones that spend least. They will be the ones that can put spend and outcome on the same page, and decide.