Vendor token forecasts missed actual usage by 25 to 50 percent, so the commitment has to be sized to floor demand
An OpenAI deal blends a seat subscription and a token meter, and only one of them behaves like software you have bought before. The commitment is set at the moment you have the least information about the metric that drives it, using a forecast produced by the party who benefits when it is high.
Prepared by Redress Compliance · August 16, 2026 · GenAI advisory. 25 to 35 enterprise GenAI procurement reviews, 2024 to 2025.
Executive summary
Vendor built token forecasts missed actual usage by 25 to 50 percent. Forecast from a pilot on your own workloads rather than from a vendor estimate, because the estimate is an opening position on volume.
Commitments sized to optimistic growth carried 10 to 20 percent shortfall risk. Size to floor demand, the consumption you would have even if adoption stalls, and buy the upside as a priced option rather than as committed spend.
Routing models by task cut token cost 15 to 30 percent with no loss of quality. Most estates send every request to a single frontier model because that is what the pilot used, and nobody revisited it before signing.
Data terms matter as much as price. Retention, training use, and processing location are set separately from commercial terms and are far harder to change after signature than a rate is.
Two pricing models in one agreement
OpenAI sells through two channels that behave nothing alike, and large deals layer committed spend and enterprise terms across both. Treating them as one negotiation is the first error.
| Channel | Metric | Behaves like | What decides the cost |
|---|---|---|---|
| ChatGPT Enterprise | Per seat | Conventional SaaS | Seat count and assignment discipline |
| API platform | Per token | Metered infrastructure | Volume, model choice, prompt and answer length |
| Committed spend | Dollar commitment | A floor, not a budget | Whether the forecast underneath it was yours |
| Enterprise terms | Contractual | Neither | Retention, training use, processing location |
Price and data handling are set separately, and the second is the harder one to change. A rate can be revisited at renewal. A retention period, a position on whether your prompts train the model, and the jurisdiction your data is processed in are architectural commitments that outlive the commercial term and are frequently signed by people evaluating the price. Read the enterprise privacy terms alongside the pricing before sizing anything.
The forecast comes from the party that benefits when it is high
Every consumption deal has the same structural problem and OpenAI deals have it in an acute form: the commitment is sized before the buyer has a run rate. Across the procurement reviews we ran, token consumption forecasts built on vendor estimates missed actual usage by 25 to 50 percent. That is not an accusation of bad faith. Nobody has reliable priors for a workload category this young, so the model is assembled from adoption assumptions rather than observation, and the party assembling it is the one whose revenue rises when the number does.
What follows is that the forecast should be treated exactly as you would treat an opening price: as a position, not as information. The correction is to forecast from a pilot on your own workloads, measuring real prompts against real documents with real users, and then to size the commitment to floor demand rather than to expected demand. Floor demand is the consumption you would still have if adoption stalled entirely, and it is the only number you can defend. Commitments sized to optimistic growth carried 10 to 20 percent shortfall risk in our file, which is money paid for volume that never arrived.
The upside should still be bought, just not as committed spend. A priced expansion option costs something to negotiate and it is almost always cheaper than committing to consumption that may not materialise, because it separates the demand you can evidence from the demand you hope for. That separation is the single most valuable structural move available in these agreements, and it is available at signature and effectively nowhere else.
Underneath the commercial question sits an engineering one worth as much. Routing models by task cut token cost 15 to 30 percent with no loss of quality across the estates reviewed. Most organisations send every request to a single frontier model because that is what the pilot used and nobody revisited the choice before signing. Classification, extraction, routing, and summarisation rarely need the heaviest model available, and matching model class to task is a day of evaluation work that recurs annually. Keeping a credible alternative model live is the same discipline viewed commercially: it is what makes the rate conversation real. The platform comparison sits in the AI platform TCO comparison, and the wider library in the GenAI practice.
- Your quote benchmarked against 500,000+ real closed deals, adjusted for size, region, and industry
- Consumption modelled from measured usage, with floor demand separated from upside
- Every risky clause flagged with the exact quote, the page, and the replacement language
The terms that outlast the rate
- Retention. How long prompts and outputs are held, and whether you can set it to zero for sensitive workloads. This is an architecture decision disguised as a contract clause.
- Training use. Whether your data trains models. Get the position in the agreement rather than from a documentation page that can be revised.
- Processing location. Where inference runs and where data rests, which decides whether the deal survives contact with your privacy function.
- Commitment shape. A floor you pay whether or not you consume, so the size matters more than the discount applied to it.
- Model deprecation. Model versions change on the vendor's schedule, so pin what you depend on or price the migration you will be forced into.
- A credible alternative. The leverage in these deals is a second model you have genuinely evaluated, not a threat you have described.
What the GenAI procurements showed, 2024 to 2025
Across roughly 25 to 35 enterprise GenAI procurement reviews, OpenAI deals were won or lost on consumption forecasting and data terms rather than on headline price:
How far token consumption forecasts built on vendor estimates diverged from actual usage.
Token cost removed by matching model class to task, with no measured loss of output quality.
Committed spend sized to optimistic growth created shortfall risk of 10 to 20 percent. The pattern is consistent: the commercial risk sits in the volume estimate rather than in the rate, and the engineering saving sits in the model choice rather than in the negotiation.
Read the vendor's own enterprise page and API pricing alongside the privacy terms before sizing a deal, because the commercial structure and the data handling position are published separately and neither implies the other.
Watch the briefing · 4:33How to Negotiate with OpenAI and Anthropic: The Vendors With Nobody to CallWhat leverage looks like against vendors with no channel, no list price, and no renewal history.
Your first five moves
- Run a pilot on your own workloads and forecast from measured tokens, treating the vendor estimate as an opening position on volume.
- Size the commitment to floor demand, the consumption that survives adoption stalling, and buy the upside as a priced expansion option.
- Route by task before you sign, matching model class to workload, since that removed 15 to 30 percent of token cost at equal quality.
- Settle retention, training use, and processing location in the agreement, not from a documentation page, because those outlast the commercial term.
- Keep a second model genuinely evaluated, because that is what makes the rate conversation real. The GenAI practice sizes the commitment with you.
Frequently asked questions
How does OpenAI sell to enterprises?
Through two channels. ChatGPT Enterprise is seat based for end users, and the API platform is consumption based for applications, with large deals layering committed spend and enterprise terms across both. They behave nothing alike and should be negotiated separately.
How accurate are vendor token forecasts?
They missed actual usage by 25 to 50 percent across the reviews we ran. That is not bad faith, since nobody has reliable priors for a workload category this young, but the model is built by the party whose revenue rises when the number does.
How should the commitment be sized?
To floor demand, meaning the consumption you would still have if adoption stalled entirely. Commitments sized to optimistic growth carried 10 to 20 percent shortfall risk. Buy the upside as a priced expansion option rather than as committed spend.
What is the cheapest saving available?
Routing models by task. It cut token cost 15 to 30 percent with no loss of quality. Most estates send every request to a single frontier model because that is what the pilot used, and nobody revisits it before signing.
Why do data terms matter as much as price?
Because a rate can be revisited at renewal and a retention period, a training use position, and a processing location cannot. They are architectural commitments that outlive the commercial term and are usually signed by people evaluating the price.
What is floor demand?
The consumption that survives a complete stall in adoption. It is the only volume number a buyer can defend from evidence, which makes it the right basis for a commitment that must be paid whether or not it is consumed.
Should we commit for the upside at all?
Not as committed spend. Negotiate a priced expansion option instead, which separates demand you can evidence from demand you hope for. That separation is available at signature and effectively nowhere else.
What happens when models are deprecated?
Versions change on the vendor schedule. Pin what you depend on in the agreement or price the migration you will eventually be forced into, because a forced model change is both a cost and a quality risk you did not budget.
Where does the leverage actually come from?
A second model you have genuinely evaluated on your own workloads. A described alternative is not leverage. An evaluated one changes what the rate conversation is about, because it supplies a comparison the vendor cannot dismiss.
Is the headline price the right focus?
No. These deals were won or lost on consumption forecasting and data terms rather than on rate. A good rate applied to a volume estimate that is 25 to 50 percent wrong produces a worse outcome than a fair rate on a defensible number.
How to Negotiate with OpenAI and Anthropic: The Vendors With Nobody to Call
Fewer than 50 sales reps globally per vendor, focused on $100M+ deals. Below $10M a negotiation rarely starts, discounts run 5 to 25 percent on commitment size, and the only leverage is credible competition between OpenAI, Anthropic, and Gemini with a benchmarked case.
Estimating the Commitment
Part 2 of the Negotiating Anthropic series. Size it on measured tokens, not on seats or headcount. How to build the baseline, how to model growth honestly, and why the error bars are wider here than in any other software category.