Skip to main content
Every rollout is led by our founder, so we take on a few at a time. Founder-led rollouts, a few at a time. Call 551-222-0087 about early pricing.
All insights

AI economics · August 3, 2026

The Token Bill Arrived. The Value Record Didn’t.

Enterprises swung from encouraging maximum AI use to clamping spend in a quarter. Both moves were made without an outcome record — and a meter cannot supply one.

The claim

Token spending is now a board-level cost line, but a metering unit is not a unit of work: without spend mapped to task classes and an observed cost per accepted outcome, neither expansion nor a clampdown is a defensible decision.

The decision

Put token spend inside the value stream: map it to named task classes with owners, observe cost per accepted unit of work, and let that record — not the size of the bill — decide what to scale, cap or stop.

The mechanism

Spend curves move for reasons unrelated to value — model mix, caching, agent delegation, provider pricing. A budget cap controls the bill but cannot say which work to cut; only an outcome record can distinguish waste from the most productive spending in the company.

The whiplash is now familiar. Leadership spends a quarter encouraging maximum AI use, then the bill lands and the same leadership orders a clampdown. Both instructions were issued without an outcome record. Encouraging use was a bet that value would follow spend; cutting it is a bet that it did not. Neither bet was ever tested, because the only number anyone could see was the meter.

A token is a metering unit, not a unit of work. It tells a finance team what was consumed, in a currency nobody has an instinct for, under accounting practices that do not yet exist for it. The bill is precise and the return is folklore — which is exactly the condition under which spending decisions swing on mood.

The meter moves for reasons that are not value

Even as a pure cost signal, the token line is unstable. Posted prices differ across providers and routes, and they track a market-share contest as much as production cost. Cached and reused context is billed at a fraction of fresh work, and agentic tools increasingly route work toward it on their own. A model-mix change, a caching improvement, or a provider reprice can move the curve sharply while the work being done — and its value — stays constant. Reading that curve as a value signal is a category error in either direction.

The common responses stay on the spend side of the ledger. Budgets and per-engineer allowances control the bill. Ratios that restate AI spending as equivalent head count make the line legible to a workforce plan. Disclosure standards, as they arrive, will make prices and usage comparable across providers. All three are useful, and none answers the operating question: which of these tokens purchased accepted work, and which purchased retries, loops and abandoned attempts? An agent that persists at a hopeless task consumes tokens at the same rate as one that ships.

That is why a clampdown is not a safe default. A cap applied to an unmapped spend line cuts the productive and the wasteful in proportion to their consumption, not their contribution. The organizations that cut confidently are not the ones spending least; they are the ones that can see cost per accepted outcome and cut where that number fails.

Put token spend inside the value stream

The decision is to change the denominator. Map token spend to named task classes — the drafting, resolution, review and build work the tokens actually attempt — each with an owner. For the material classes, observe cost per accepted unit of work: the attempt that survived review and was used, not the attempt that ran. Then let that record make the spending decisions a bill cannot: raise the cap where unit cost is strong and demand is real, reroute where a cheaper path meets the acceptance threshold, and stop the classes where spend accumulates without accepted output.

Keep the record honest about what each number is. Cash is a documented and attributable change in money paid or received — a cancelled licence, a reduced external-services invoice, a lower unit cost on the same accepted volume. Released people-time is capacity and stays in operational units. Equivalent-head-count ratios and share-of-operating-expense scenarios are modeled context: useful for planning, never additive to a value total. These categories belong beside one another; they do not form a defensible sum.

Date the record. Token prices, model capability and caching behavior all move; a unit cost observed in one quarter is an assumption by the next. Every task-class entry needs a review date, and a material price or capability change should pull that review forward.

What would count as proof?

  • Token spend mapped to named task classes with owners, covering the material share of the bill — not a sample.
  • Cost per accepted unit of work observed on the workflow for each material class, with the acceptance threshold stated.
  • At least one dated spending decision — a cap raised or lowered, a route switched, a workload stopped — made because of the record, with the before and after retained.
  • Cash effects documented and attributed where claimed; capacity kept in operational units; modeled ratios and scenarios labelled as context.
  • A review date on every task-class entry, with evidence that a price or capability change triggered an early review at least once.

Each item is checkable from the record itself. None requires trusting a narrative about what the spending probably achieved.

What remains unclaimed?

A falling token bill is not established savings, and a rising one is not established investment; both require the outcome record to mean anything. An equivalent-head-count ratio is not a productivity result. A cheaper route is not value until the same acceptance threshold is met on the same work. And no spend-side total — however precisely metered — becomes a value claim by being large, disciplined or benchmarked. Value enters the record through accepted work and documented cash, or it remains modeled upside.

The bill becomes decidable when it reaches accepted work

From token bill to accepted work. The bill is one number. Decisions need the path from spend through task classes to an observed unit cost with a review date.
A left-to-right waterfall decomposes one Token bill into named Task class columns, each flowing to Accepted work and an observed Unit cost, closing at a Review date. Spend that never reaches accepted work is visible as the gap between the bill and the accepted columns, and that gap — not the size of the bill — is what a spending decision should act on.

One number invites a mood. A mapped record invites a decision. The difference is whether the bill can be traced to work someone accepted.

Where to take this next

Read the Oabo method

Related stories

July 17, 2026

Smarter Models Make Yesterday’s Value Case Expire.