Laptop screen showing charts and financial figures
Enterprise AI Costs

AI cost management in 2026. Three meters, each with its own control.

How enterprise AI spend splits across tokens, assistant seats and cloud hosted models, with worked numbers and the contract terms that keep each meter in check.

Contact Us GenAI Advisory
500+Enterprise clients
$2B+Under advisory
PublishedDecember 15, 2025UpdatedSeptember 25, 2026
ContentsKey takeawaysThe three AI metersWhy token forecasts missSavings from model routingReclaiming assistant seatsFinding shadow AI spendSizing an AI commitmentWhat our reviews showedWhat to do nextFAQ

Enterprise AI spend runs on three meters with different economics: API tokens, per seat assistants and cloud hosted models. Governance built for any one of them fails on the other two, so each needs its own owner and control.

Key takeaways
  • Three meters. API tokens, per seat assistants and cloud hosted models share a budget line, and each needs its own control and owner.
  • Forecasts miss both ways. Token forecasts missed actuals by 2 to 4 times in the first year, so measurement protects you better than caution.
  • Routing is the largest saving. Moving eligible workloads to smaller models cut token costs 30 to 60 percent at acceptable quality in most tested cases.
  • Seats sit idle. Assistant seats in active use ran at 40 to 60 percent of licensed seats, and ordinary SaaS reclamation fixes it.
  • Commit to measured volume. A commitment transfers forecast risk to you, so size it on consumption you have measured and ask for rollover and model substitution.
  • Shadow AI is already spend. Team subscriptions small enough to expense accumulate before procurement ever sees a line item.

What does AI cost management cover in an enterprise?

AI cost management covers three separate meters: API token usage, per seat assistant subscriptions, and cloud hosted model consumption. They share a budget line and nothing else. Each one bills on its own logic, fails in its own way and needs its own owner.

The three meters of enterprise AI spend
MeterBilling logicPrimary controlHow it failsUsual owner
API tokensPer token consumed, by model classModel routing and prompt designForecast error in both directionsEngineering or platform team
Per seat assistantsPer licensed user, renewing like SaaSAssignment discipline and reclamationDormant seats and low active useSoftware asset management
Cloud hosted modelsMetered through the cloud platformCommitment sizing against a measured baselineCommitted spend locking forecast risk onto youCloud FinOps with procurement

A control built for one meter has no effect on the other two. Seat governance, which most organizations already run well, does nothing to token spend. Routing is an engineering discipline and leaves dormant assistant licenses untouched. Commitment sizing is a procurement exercise and addresses neither.

Split the budget line by meter and give each one a named owner who reports on its trend every month. Our enterprise AI governance guide covers the policy side: approved tools, data rules and who signs off on new use cases.

Watch the briefingEpisode 2 of 6 · 3:52

Why do AI token forecasts miss so badly?

In the first year of production workloads, token spend forecasts missed actuals by 2 to 4 times, and they missed in both directions. Some organizations overbought and stranded the unused balance. Others underbought and paid overage rates on the excess.

Both outcomes had one cause. No team has a reliable prior for how a generative workload behaves once real users reach it, and a pilot with a few friendly testers predicts little. This is a different problem from a vendor talking a buyer into an oversized commitment.

Why a cautious commitment does not fix a two way error

If forecasts only ever ran high, committing less than the vendor proposes would be enough. Because they run wrong in either direction, caution protects you against one error and exposes you to the other. The better protection is measurement, set up in three steps.

  • Forecast from your own consumption data. Use tokens actually billed on your workloads, split by model and application. Treat vendor projections as a sales document.
  • Size committed spend against a measured baseline. Commit to what you already consume and let growth run on demand until it shows up in the data.
  • Revisit both quarterly through the first year. That is when the adoption curve is steepest and least predictable, and an annual cycle finds the error a full year after it started compounding.

Worked example: what each kind of miss costs

Say a vendor projects $600,000 of token usage at list price for year one and offers a discount for committing to it. Assume a 10 percent discount for illustration, so you commit $540,000. The table shows three outcomes, with usage above the commitment billed at list.

Hypothetical $540,000 commitment against three actual outcomes
Actual usage at listCommitment consumedBilled above commitmentTotal paidResult
Half the forecast, $300,000$270,000$0$540,000$270,000 stranded; you paid 80 percent above list for what you used
On forecast, $600,000$540,000$0$540,000Discount earned in full, worth $60,000
Double the forecast, $1,200,000$540,000$600,000$1,140,000$600,000 over the approved budget

Measured as waste, the overbuy is worse: $270,000 paid for nothing, against $60,000 of discount forgone on the excess in the underbuy. The underbuy still breaks the approved budget by $600,000, and where overage bills at a premium to list the gap narrows. In every row the forecast error decided the result far more than the discount did.

Free white paper

Enterprise AI Procurement Strategy Brief

Commitment sizing, renewal timing and data terms for AI platform contracts.

Get the white paper →

How much does model routing cut AI token costs?

Routing eligible workloads to smaller models cut token costs by 30 to 60 percent in our reviews, with acceptable quality in most tested cases. It is the largest single saving on the token meter. It beats any discount available on these agreements, needs no negotiation, and repeats every year as model options change.

Routing means sending each type of request to the cheapest model that passes your own acceptance tests. Classification, extraction and ticket triage rarely need a frontier model.

Worked example: moving half the traffic to a small model

Take a hypothetical workload of 10,000 million input tokens and 1,000 million output tokens a month. Use illustrative rates of $3 input and $15 output per million tokens on the large model, and $0.25 and $1.25 on the small one. Check current rate cards before you reuse these figures.

Monthly cost of one workload, before and after routing
OptionInput costOutput costMonthly total
Everything on the large model10,000 × $3 = $30,0001,000 × $15 = $15,000$45,000
Half stays on the large model5,000 × $3 = $15,000500 × $15 = $7,500$22,500
Half routed to the small model5,000 × $0.25 = $1,250500 × $1.25 = $625$1,875
Routed total$24,375

The routed workload costs $20,625 a month less, a cut of about 46 percent, or $247,500 a year. The saving lands a little below the share of traffic routed, because the small model still costs something.

Why routing does not happen on its own

The obstacle is organizational. The engineer choosing the model is solving a quality problem under time pressure and has no view of the marginal cost. The person holding the budget has no view of the model choice.

  • Give engineers the cost per request. Show it next to latency and accuracy in the evaluation results for every candidate model.
  • Give finance the model mix. Report spend by model and application monthly, so a workload running entirely on the largest model stands out.
  • Keep an acceptance test set per workload. Rerun it whenever a new small model ships.
  • Use batch pricing for work that can wait. The OpenAI Batch API and the Anthropic Message Batches API both bill at 50 percent of standard prices, for jobs processed within a 24 hour window.
Developer working at a desk with monitoring dashboards on several screens
Most model choices are made in an evaluation run weeks before finance sees an invoice, so cost per request has to appear in the evaluation itself.

Why we would not open with the token rate discount

The usual advice is to open an AI negotiation by pushing for the deepest discount on the committed token rate. We think that gets the order wrong. A rate agreed before routing prices traffic that is about to move to cheaper models, and deeper discounts usually need a larger commitment, where the two way forecast error does its damage.

Route and measure first. Then negotiate the rate on the smaller, measured volume, with the contract terms set out below.

How do you cut wasted AI assistant seats?

Run the same reclamation process you already run for any SaaS product. Per user AI assistants renew like ordinary SaaS and carry the same dormant seat problem. In the organizations we measured, seats in active use ran at 40 to 60 percent of licensed seats.

The seat meter is the easiest of the three and the most often ignored, because it looks solved. It is also the cheapest win, since the process already exists and only needs pointing at a new line item.

How to check who actually uses the seats

  • Microsoft 365 Copilot. In the Microsoft 365 admin center, open Reports, then Usage, then Microsoft Copilot. The report compares enabled users with active users over 7, 28, 90 or 180 days and lists each user's last activity date by app. Our note on Copilot monthly active users covers how to read it.
  • ChatGPT Enterprise. Workspace analytics in the admin console shows member activity and can be exported.
  • Your identity provider. Single sign on logs show who has never opened a tool at all, which covers assistants whose own reporting is thin.

Worked example: reclaiming a 2,000 seat assistant deal

Say you license 2,000 assistant seats at a hypothetical $30 per user per month, or $720,000 a year. A 90 day review finds 1,000 users active, which is 50 percent. You keep the active users plus a 10 percent buffer for joiners, so 1,100 seats.

Cutting 900 seats removes $324,000 a year at that price (900 × $30 × 12). Annual seat subscriptions usually allow reductions only at renewal or anniversary, so the review has to finish before the notice date, or the saving slips a full year.

How do you find shadow AI spend?

Look where small purchases land: expense claims, corporate card statements and sign in logs. Shadow AI is real spend before procurement sees it. Team subscriptions accumulate one at a time, each small enough to expense, so none of them crosses the threshold a budget review looks at.

  1. Search expense and card data. Filter the past year by AI vendor merchant names, including API providers that engineering teams pay by card during a pilot.
  2. Check sign in and network logs. Look for traffic to AI tools that are not on the approved list.
  3. Move real users onto the enterprise agreement. Weekly users belong under your negotiated price and data terms. Close the rest, and have finance reject AI subscriptions on expense claims once an approved tool exists.

Our shadow AI spend report covers where these purchases usually hide.

How should you size an AI commitment for cloud hosted models?

Size it against a measured baseline of what you already consume, never against the adoption plan. Hosted models on Azure, AWS and Google Cloud are metered through the cloud platform, on meters of their own. A commitment buys a discount and in exchange transfers the forecast risk to you.

That is a poor trade when the forecast is the least reliable number in the deal. Set the commitment at recent measured consumption, with a ramp only where growth is already visible. The enterprise AI platform TCO comparison shows how the platforms price the same workload.

Contract terms to ask for

  • Quarterly resizing in year one. It matches the contract to the period when forecasts are least reliable.
  • Rollover of unused balance. Unused commitment carries into the next term instead of expiring, which caps the cost of overbuying.
  • Overage at the committed rate. Usage above the commitment bills at the discounted rate with no premium, which caps the cost of underbuying.
  • Model substitution. The commitment and rate apply to any model the vendor offers, including new releases, so routing never strands the balance. See the swap and reallocation rights clause.
  • Drawdown of cloud commitments. Confirm in writing whether AI consumption counts against your Azure, AWS or Google Cloud commitment (see cloud AI commitment negotiation).

What the account team will say, and what to answer

Typical vendor lines on AI commitments and seat deals
What you will hearWhat to say back
"Commit to the three year volume and we can give you our best rate.""We will commit to measured consumption with quarterly resizing. Price the rate on that."
"Unused commitment cannot roll over.""Then the commitment stays at our measured baseline, and usage above it bills at the committed rate."
"The discount applies to this model family only.""Apply it to total spend, whichever model the traffic runs on, including releases during the term."
"Seat counts can only change at renewal.""Then we buy the active seat count now and add seats each quarter as usage proves out."

Where to find your measured baseline

  • OpenAI. The Usage and Costs endpoints of the Admin API, by project, API key and model.
  • Anthropic. The Usage and Cost Admin API, by model, workspace, API key and service tier.
  • Microsoft Azure. Microsoft Cost Management, filtered on your Azure OpenAI and Foundry resources.
  • AWS. Cost Explorer for Amazon Bedrock, with application inference profiles and cost allocation tags to split spend by application.
  • Google Cloud. The Cloud Billing export to BigQuery, filtered on Vertex AI services.

What did our AI cost reviews show in 2024 and 2025?

Across roughly 20 to 30 enterprise AI contract and cost reviews in 2024 and 2025, AI was the fastest growing and least governed line in the software budget. The ranges on this page come from those reviews, and three patterns repeated.

  1. Blended reporting hid the meter that was growing. Public rate cards set the reference token prices and cloud platforms meter hosted models separately, so one AI total could not show which meter drove the increase.
  2. Both forecast errors appeared in the same period. Some organizations stranded commitment while others paid overage, and the cause was the same: no reliable prior for how a workload behaves once real users reach it.
  3. Organizations governed the meter they already knew. Where AI was treated as one budget line, the team applied whichever control it was already good at and left the other two meters ungoverned.
A commitment buys a discount and hands you the forecast risk. Sign one only against numbers you have measured on your own workloads.

What to do next

  1. This month. Split the AI budget into its three meters and assign a named owner and control to each, instead of governing the total.
  2. Within 90 days. Measure token consumption on your own workloads by model and application, and rebuild the forecast from it, expecting the first pass to be wrong in either direction.
  3. Next quarter. Run a routing exercise across live workloads, with an acceptance test set for each one.
  4. Before the seat renewal notice date. Reclaim dormant assistant seats with the SaaS reclamation process you already run.
  5. Before you sign any commitment. Size it against the measured baseline, ask for rollover, overage at the committed rate and model substitution, and close the shadow subscription routes. Our GenAI practice builds the cost model with you.
  6. Every quarter in year one. Revisit the forecast and the commitment together. Vendor guides sit in the GenAI knowledge hub.

Frequently asked questions

What are the three meters of enterprise AI spend?

API token usage billed by model, per seat assistant subscriptions billed per user, and hosted model consumption billed through Azure, AWS or Google Cloud. One model family can reach you on all three at once.

How wrong are token forecasts?

Wrong by a multiple in the first year of production, sometimes too high and sometimes too low. Prompts and retrieved context grow after launch, retries add volume, and adoption spreads past the pilot group, none of which pilot traffic captures.

Why does the both directions detail matter?

It removes caution as a strategy, since a small commitment only protects you when the forecast is too high. Protection comes from terms that make either miss cheap: rollover for overbuying and overage at the committed rate for underbuying.

What is the largest single saving?

Model routing on the token meter. Batch processing for work that can wait and prompt caching for long repeated instructions add to it, and none of these steps needs vendor approval or a contract change.

Why does routing not happen by default?

Prototypes are built on the most capable model to get quality working, and that choice survives into production. Make the model mix a standing line in the monthly AI cost review so someone revisits it.

How bad is assistant seat waste?

Roughly half of licensed seats sat idle in the organizations we measured. Start the review with seats assigned in bulk by department during the first rollout, since those were never matched to a named request.

Should we take a committed spend discount?

Yes, when it is sized to consumption you have measured and unused balance can roll over. Work out what you would strand if usage came in at half the forecast. If that exceeds the discount, commit less.

What is shadow AI costing?

Usually more than finance expects, and the figure is only known after a search of expense, card and sign in data. Team plans also sit outside your enterprise data terms and outside the volume that sets your price.

Can one control cover all three meters?

No. Tokens respond to routing, seats to reclamation and cloud hosted models to commitment sizing. The meters can share one monthly report, with a named owner answerable for each trend.

How often should the forecast be revisited?

Every quarter through the first year of production, then twice a year once consumption settles. Time each review so an updated forecast is ready before any true up or seat renewal notice deadline.

Newsletter
Licensing news that changes what you pay

One email a week on vendor price moves, audit activity and what worked in recent renewals.

Subscribe
Vendor Shield
An advisor on call for every vendor conversation

Always on advisory for renewals, audits and contract questions across your software vendors.

Explore Vendor Shield
Advisory White Paper

Get the enterprise AI procurement strategy brief.

Sourcing, contracting and renewal across AI platforms, with the commitment arithmetic and the data terms that decide whether a deal is safe to sign.

Gated with a work email on the download page. No sales follow up you did not ask for.

Get the White Paper →
We never share your details with vendors.

AI platform licensing news, once a week.

Price changes, audit activity and what worked in recent renewals. No vendor spin.