Token forecasts missed actuals by 2 to 4 times in both directions, and each of the three meters needs its own control
AI spend is not one line with one control. It is three meters with different economics, and a governance model built for any one of them fails on the other two.
Prepared by Redress Compliance · August 16, 2026 · GenAI advisory. 20 to 30 enterprise AI contract and cost reviews, 2024 to 2025.
Executive summary
Token forecasts missed actuals by 2 to 4 times in the first year, in both directions. That symmetry matters: over forecasting wastes committed spend and under forecasting triggers overage, so accuracy is worth more than either optimism or caution.
Assistant seat utilisation ran at 40 to 60 percent of licensed seats. Per user AI assistants renew like ordinary SaaS and carry the same dormant seat problem, which means the oldest lever in software applies directly.
Routing eligible workloads to smaller models cut token costs 30 to 60 percent with acceptable quality in most tested cases. It is the largest single lever on the token meter and it requires no negotiation.
Shadow AI is real spend before procurement sees it. Ungoverned team level subscriptions accumulate, and they are invisible in a budget review because each one is small enough to expense.
Three meters, three different controls
Enterprise AI spend arrives as API token usage, per seat assistant subscriptions, and cloud hosted model consumption. They share a budget line and nothing else.
| Meter | Billing logic | Primary control | How it fails |
|---|---|---|---|
| API tokens | Per token consumed, by model class | Model routing and prompt design | Forecast error in both directions |
| Per seat assistants | Per licensed user, renewing like SaaS | Assignment discipline and reclamation | Dormant seats at 40 to 60 percent utilisation |
| Cloud hosted models | Metered through the cloud platform | Commitment sizing against measured baseline | Committed spend locking forecast risk onto you |
A control built for one meter does nothing for the other two. Seat governance, which most organisations already run well, has no effect on token spend. Token routing, which is an engineering discipline, does nothing about dormant assistant licences. And commitment sizing, which is a procurement exercise, addresses neither. Estates that treated AI as a single budget line applied whichever control they were already good at and left the other two ungoverned.
The forecast misses in both directions, which changes what to optimise for
Token spend forecasts missed actuals by 2 to 4 times in the first year of production workloads, and the detail that matters most is that they missed in both directions. This is not the familiar story of a vendor talking a buyer into an oversized commitment. Some estates dramatically overbought and stranded the balance; others dramatically underbought and paid overage rates on the excess. Both outcomes came from the same cause, which is that nobody has a reliable prior for how a generative workload behaves once real users reach it.
That symmetry changes the objective. If forecasts only ever ran high, the correct posture would be caution, and buyers could simply commit less than proposed. Because they run wrong in both directions, caution is not a strategy either, and the thing worth investing in is measurement rather than conservatism. Forecast from consumption data on your own workloads rather than from vendor projections, size committed spend against a measured baseline rather than against ambition, and revisit both quarterly rather than annually, because the first year is when the curve is steepest and least predictable.
The largest available saving sits on the engineering side rather than the commercial one. Routing eligible workloads to smaller models cut token costs by 30 to 60 percent with acceptable quality in most tested cases. That is a bigger number than any discount available on these agreements, it requires no negotiation, and it is repeatable annually as model options change. What stops it happening is organisational rather than technical: the person choosing the model is solving a quality problem under time pressure and has no view of the marginal cost, while the person holding the budget has no view of the model choice.
The seat meter is the easiest of the three and the most often ignored, because it looks solved. Per user AI assistants renew like ordinary SaaS and carry the same dormant seat problem, running at 40 to 60 percent utilisation of licensed seats. Every reclamation discipline an organisation already applies to conventional software applies here unchanged, which makes it the cheapest of the three wins. Shadow AI then sits underneath all of it as ungoverned team level subscriptions that accumulate below the threshold anyone reviews. The governance framing sits in the AI governance playbook, the platform comparison in the platform TCO comparison, and the wider library in the GenAI practice.
- Your quote benchmarked against 500,000+ real closed deals, adjusted for size, region, and industry
- Consumption modelled per meter, with committed spend sized against a measured baseline
- Every risky clause flagged with the exact quote, the page, and the replacement language
The control per meter
- Tokens: route by task before you negotiate rate. Smaller models cut cost 30 to 60 percent at acceptable quality in most tested cases, which exceeds any discount available on the agreement.
- Tokens: forecast from your own consumption data, not from vendor projections, and revisit quarterly through the first year when the curve is steepest.
- Seats: run reclamation exactly as you would for any SaaS, since utilisation at 40 to 60 percent is the same dormant seat problem with a new label.
- Cloud hosted: size commitments against a measured baseline, never against ambition, because a commit buys a discount and transfers the forecast risk to you.
- All three: close the shadow routes, because ungoverned team subscriptions accumulate below the threshold anybody reviews.
- Do not apply one control to all three meters. Estates that did applied whichever discipline they already had and left the rest ungoverned.
What the AI cost reviews showed, 2024 to 2025
Across roughly 20 to 30 enterprise AI contract and cost reviews, AI spend was the fastest growing and least governed line in the software budget:
How far token spend forecasts missed actuals in the first year of production workloads, in both directions.
Token cost removed by routing eligible workloads to smaller models, at acceptable quality in most tested cases.
AI assistant seat utilisation ran at 40 to 60 percent of licensed seats in the estates measured. That is the oldest problem in software licensing arriving on a new line item, and the existing reclamation discipline solves it without modification.
Public rate cards set the token anchors and cloud platforms meter hosted models separately, so a single blended view of AI spend hides which meter is actually moving.
Watch the briefing · 4:33From Licenses to Subscriptions to AI Consumption: The Third Repricing of SoftwareWhy AI spend behaves unlike the two software pricing eras that came before it.
Your first five moves
- Split the AI budget into its three meters and assign a named control to each rather than governing the total.
- Measure token consumption on your own workloads and rebuild the forecast from it, expecting to be wrong in either direction on the first pass.
- Run a model routing exercise across live workloads, since it is worth more than any discount on the contract and needs no negotiation.
- Reclaim dormant assistant seats using the SaaS discipline you already run, against 40 to 60 percent measured utilisation.
- Size any committed spend against the measured baseline and close the shadow subscription routes. The GenAI practice builds the model with you.
Frequently asked questions
What are the three meters of enterprise AI spend?
API token usage, per seat assistant subscriptions, and cloud hosted model consumption. They share a budget line and nothing else, and a control built for one does nothing for the other two.
How wrong are token forecasts?
They missed actuals by 2 to 4 times in the first year of production workloads, in both directions. Some estates overbought and stranded the balance, others underbought and paid overage, both from the same cause.
Why does the both directions detail matter?
Because it rules out caution as a strategy. If forecasts only ran high you could simply commit less. Since they run wrong either way, the thing worth investing in is measurement rather than conservatism.
What is the largest single saving?
Model routing. Sending eligible workloads to smaller models cut token costs 30 to 60 percent at acceptable quality in most tested cases, which is larger than any discount available on these agreements and requires no negotiation.
Why does routing not happen by default?
Because the person choosing the model is solving a quality problem under time pressure with no view of marginal cost, while the person holding the budget has no view of the model choice. It is an organisational gap rather than a technical one.
How bad is assistant seat waste?
Utilisation ran at 40 to 60 percent of licensed seats. Per user AI assistants renew like ordinary SaaS and carry the same dormant seat problem, so the reclamation discipline you already run applies unchanged.
Should we take a committed spend discount?
Only sized against a measured baseline. A commitment buys a discount and transfers the forecast risk onto you, which is a poor trade when the forecast is the least reliable number in the deal.
What is shadow AI costing?
More than it appears, because each subscription is small enough to expense and therefore sits below the threshold anyone reviews. It accumulates as real spend before procurement ever sees a line item.
Can one control cover all three meters?
No, and estates that tried applied whichever discipline they were already good at and left the other two ungoverned. Seat governance does nothing for tokens, routing does nothing for dormant licences, and commitment sizing addresses neither.
How often should the forecast be revisited?
Quarterly through the first year, not annually. That is when the adoption curve is steepest and least predictable, and an annual cycle discovers the error a full year after it started compounding.
Estimating the Commitment
Part 2 of the Negotiating Anthropic series. Size it on measured tokens, not on seats or headcount. How to build the baseline, how to model growth honestly, and why the error bars are wider here than in any other software category.