HomeGenAI HubAI Cost Management
GenAI  |  Cost Management Cost Control Playbook 2026

Token forecasts missed actuals by 2 to 4 times in both directions, and each of the three meters needs its own control

AI spend is not one line with one control. It is three meters with different economics, and a governance model built for any one of them fails on the other two.

Prepared by Redress Compliance · August 16, 2026 · GenAI advisory. 20 to 30 enterprise AI contract and cost reviews, 2024 to 2025.

Executive summary

Token forecasts missed actuals by 2 to 4 times in the first year, in both directions. That symmetry matters: over forecasting wastes committed spend and under forecasting triggers overage, so accuracy is worth more than either optimism or caution.

Assistant seat utilisation ran at 40 to 60 percent of licensed seats. Per user AI assistants renew like ordinary SaaS and carry the same dormant seat problem, which means the oldest lever in software applies directly.

Routing eligible workloads to smaller models cut token costs 30 to 60 percent with acceptable quality in most tested cases. It is the largest single lever on the token meter and it requires no negotiation.

Shadow AI is real spend before procurement sees it. Ungoverned team level subscriptions accumulate, and they are invisible in a budget review because each one is small enough to expense.

2 to 4x
How far token forecasts missed actuals, in both directions.
30 to 60%
Token cost cut by routing eligible workloads to smaller models.
40 to 60%
Assistant seat utilisation against licensed seats.
3
Meters carrying enterprise AI spend, each needing its own control.
1.

Three meters, three different controls

Enterprise AI spend arrives as API token usage, per seat assistant subscriptions, and cloud hosted model consumption. They share a budget line and nothing else.

MeterBilling logicPrimary controlHow it fails
API tokensPer token consumed, by model classModel routing and prompt designForecast error in both directions
Per seat assistantsPer licensed user, renewing like SaaSAssignment discipline and reclamationDormant seats at 40 to 60 percent utilisation
Cloud hosted modelsMetered through the cloud platformCommitment sizing against measured baselineCommitted spend locking forecast risk onto you

A control built for one meter does nothing for the other two. Seat governance, which most organisations already run well, has no effect on token spend. Token routing, which is an engineering discipline, does nothing about dormant assistant licences. And commitment sizing, which is a procurement exercise, addresses neither. Estates that treated AI as a single budget line applied whichever control they were already good at and left the other two ungoverned.

2.

The forecast misses in both directions, which changes what to optimise for

Token spend forecasts missed actuals by 2 to 4 times in the first year of production workloads, and the detail that matters most is that they missed in both directions. This is not the familiar story of a vendor talking a buyer into an oversized commitment. Some estates dramatically overbought and stranded the balance; others dramatically underbought and paid overage rates on the excess. Both outcomes came from the same cause, which is that nobody has a reliable prior for how a generative workload behaves once real users reach it.

That symmetry changes the objective. If forecasts only ever ran high, the correct posture would be caution, and buyers could simply commit less than proposed. Because they run wrong in both directions, caution is not a strategy either, and the thing worth investing in is measurement rather than conservatism. Forecast from consumption data on your own workloads rather than from vendor projections, size committed spend against a measured baseline rather than against ambition, and revisit both quarterly rather than annually, because the first year is when the curve is steepest and least predictable.

The largest available saving sits on the engineering side rather than the commercial one. Routing eligible workloads to smaller models cut token costs by 30 to 60 percent with acceptable quality in most tested cases. That is a bigger number than any discount available on these agreements, it requires no negotiation, and it is repeatable annually as model options change. What stops it happening is organisational rather than technical: the person choosing the model is solving a quality problem under time pressure and has no view of the marginal cost, while the person holding the budget has no view of the model choice.

The seat meter is the easiest of the three and the most often ignored, because it looks solved. Per user AI assistants renew like ordinary SaaS and carry the same dormant seat problem, running at 40 to 60 percent utilisation of licensed seats. Every reclamation discipline an organisation already applies to conventional software applies here unchanged, which makes it the cheapest of the three wins. Shadow AI then sits underneath all of it as ungoverned team level subscriptions that accumulate below the threshold anyone reviews. The governance framing sits in the AI governance playbook, the platform comparison in the platform TCO comparison, and the wider library in the GenAI practice.

Try Vera AI · free 30 day trial
Vera splits AI spend by meter and rebuilds each forecast from measured consumption.
  • Your quote benchmarked against 500,000+ real closed deals, adjusted for size, region, and industry
  • Consumption modelled per meter, with committed spend sized against a measured baseline
  • Every risky clause flagged with the exact quote, the page, and the replacement language
Start the free Vera AI trial →30 days free · no credit card · cancel anytime
3.

The control per meter

4.

What the AI cost reviews showed, 2024 to 2025

Across roughly 20 to 30 enterprise AI contract and cost reviews, AI spend was the fastest growing and least governed line in the software budget:

2 to 4x
Forecast error

How far token spend forecasts missed actuals in the first year of production workloads, in both directions.

30 to 60%
The routing saving

Token cost removed by routing eligible workloads to smaller models, at acceptable quality in most tested cases.

AI assistant seat utilisation ran at 40 to 60 percent of licensed seats in the estates measured. That is the oldest problem in software licensing arriving on a new line item, and the existing reclamation discipline solves it without modification.

Public rate cards set the token anchors and cloud platforms meter hosted models separately, so a single blended view of AI spend hides which meter is actually moving.

Watch the briefing · 4:33From Licenses to Subscriptions to AI Consumption: The Third Repricing of SoftwareWhy AI spend behaves unlike the two software pricing eras that came before it.
5.

Your first five moves

  1. Split the AI budget into its three meters and assign a named control to each rather than governing the total.
  2. Measure token consumption on your own workloads and rebuild the forecast from it, expecting to be wrong in either direction on the first pass.
  3. Run a model routing exercise across live workloads, since it is worth more than any discount on the contract and needs no negotiation.
  4. Reclaim dormant assistant seats using the SaaS discipline you already run, against 40 to 60 percent measured utilisation.
  5. Size any committed spend against the measured baseline and close the shadow subscription routes. The GenAI practice builds the model with you.
6.

Frequently asked questions

What are the three meters of enterprise AI spend?

API token usage, per seat assistant subscriptions, and cloud hosted model consumption. They share a budget line and nothing else, and a control built for one does nothing for the other two.

How wrong are token forecasts?

They missed actuals by 2 to 4 times in the first year of production workloads, in both directions. Some estates overbought and stranded the balance, others underbought and paid overage, both from the same cause.

Why does the both directions detail matter?

Because it rules out caution as a strategy. If forecasts only ran high you could simply commit less. Since they run wrong either way, the thing worth investing in is measurement rather than conservatism.

What is the largest single saving?

Model routing. Sending eligible workloads to smaller models cut token costs 30 to 60 percent at acceptable quality in most tested cases, which is larger than any discount available on these agreements and requires no negotiation.

Why does routing not happen by default?

Because the person choosing the model is solving a quality problem under time pressure with no view of marginal cost, while the person holding the budget has no view of the model choice. It is an organisational gap rather than a technical one.

How bad is assistant seat waste?

Utilisation ran at 40 to 60 percent of licensed seats. Per user AI assistants renew like ordinary SaaS and carry the same dormant seat problem, so the reclamation discipline you already run applies unchanged.

Should we take a committed spend discount?

Only sized against a measured baseline. A commitment buys a discount and transfers the forecast risk onto you, which is a poor trade when the forecast is the least reliable number in the deal.

What is shadow AI costing?

More than it appears, because each subscription is small enough to expense and therefore sits below the threshold anyone reviews. It accumulates as real spend before procurement ever sees a line item.

Can one control cover all three meters?

No, and estates that tried applied whichever discipline they were already good at and left the other two ungoverned. Seat governance does nothing for tokens, routing does nothing for dormant licences, and commitment sizing addresses neither.

How often should the forecast be revisited?

Quarterly through the first year, not annually. That is when the adoption curve is steepest and least predictable, and an annual cycle discovers the error a full year after it started compounding.

Watch the briefingEpisode 2 of 6 · 3:52

Estimating the Commitment

Part 2 of the Negotiating Anthropic series. Size it on measured tokens, not on seats or headcount. How to build the baseline, how to model growth honestly, and why the error bars are wider here than in any other software category.

© 2026 Redress Compliance · Independent, buyer sideredresscompliance.com
Industry Recognized
500+ Enterprise Clients
$2B+ Under Advisory
11 Vendor Practices
100% Buyer Side Independent
GenAI White Paper

The full enterprise AI procurement strategy brief from the GenAI practice.

Sourcing, contracting, and renewal across the AI platform layer, with the commitment arithmetic and the data terms that decide whether a deal is safe to sign.

Gated with a work email on the download page. No sales follow up you did not ask for.

Get the White Paper →
Independent, buyer side. We never share your details with vendors.
Run the software spend health check against your AI estate in under five minutes.
Open the Tool → GenAI Practice →
Editorial boardroom interior

The advisor your vendors do not want.

500+ enterprise clients. 11 vendor practices. Industry recognized. One conversation can change what you pay for the next three years.

Stay ahead of AI platform pricing and contract moves.

One buyer side briefing a week. Renewal signals, discount bands, and the levers that work. No vendor spin.