Contents
Key takeawaysThe four cost driversWhat we saw in recent decisionsPlatforms comparedPilot versus productionWorked TCO exampleIntegration and governanceSizing the commitmentPricing the exitWhat to do nextFAQToken price is usually the smallest line in an enterprise AI budget at scale. Compare platforms on production inference, commitment terms, integration effort and exit cost for your real workload, and the ranking often changes.
- Four drivers set the total. Inference at production volume, commitment terms, integration and data engineering, and exit cost decide what an AI platform costs, and the token rate feeds only the first.
- The pilot is a floor. Production inference ran 3 to 6 times the pilot that justified the platform choice in the decisions we advised.
- Integration is the year one expense. Connecting the model to your data and systems was the largest first year cost that never shows on a rate card.
- Commit to the low forecast. Committed spend lowers the rate and forfeits what you do not use, so size it to conservative demand and let growth run at standard rates.
- Price the exit before you sign. Reworking prompts, evaluations and pipelines for another model is real engineering work, and it can reverse a year one ranking.
- Governance recurs. Monitoring, evaluation and compliance review are annual engineering costs that grow with each use case put live.
Most enterprise AI platform shortlists compare one number, the price per million tokens. At production scale that is usually the smallest line in the budget. The total is set by four drivers: inference volume at production scale, commitment terms, integration and data engineering effort, and the cost of leaving.
This page shows how to compare a direct model API with Azure OpenAI, Amazon Bedrock and Google's Vertex AI on all four, with worked examples you can copy into your own model. It draws on the 30 enterprise AI platform decisions we advised across 2024 and 2025.
What makes up the total cost of an enterprise AI platform?
Four drivers make up the total: inference at production volume, the commitment you sign, the integration and data work, and exit cost. The token rate is one input to the first of them. Governance sits alongside integration as a recurring engineering cost.
| Cost line | What it covers | When you pay | How to estimate it |
|---|---|---|---|
| Inference at production scale | Request volume, token length per request, peak concurrency | Monthly invoice | Measured tokens per request times a signed off production request count |
| Integration and data engineering | Connecting the model to your data, identity, applications and logging | Mostly year one | Engineering days per data source and per application, priced at your internal rate |
| Governance | Monitoring, evaluation and compliance review | Every year | A fixed effort per live use case, repeated at each model version change |
| Commitment | Committed spend or reserved capacity | Up front or monthly, forfeited if unused | Discount earned minus the value of committed usage you do not consume |
| Exit | Reworking prompts, evaluations and pipelines for another model | When you switch | Rework, retesting and parallel running for each use case |
Integration and governance together often outweigh the gap between two rate cards, as the worked example below shows. That is why a platform that fits cheaply into the cloud you already run can beat one with a lower unit rate.
What You Are Actually Buying
What have we seen in recent enterprise AI platform decisions?
Across the platform decisions we advised from 2024 to 2025, the token rate was rarely what decided the total once all four drivers were modeled. Three patterns came up again and again.
- Production outgrew the pilot. Inference at production scale cost 3 to 6 times the pilot that justified the platform choice, and the typical scale up was about four times.
- Integration was the real year one expense. Integration and data engineering added between 20 and 50 percent on top of platform cost in the first year, with a median of 35 percent.
- Exit was left out. No buyer in the file had priced a switch before signing, so later switches took longer and cost more than expected.
Why we do not pick the platform with the lowest token price
The usual advice is to choose the lowest price per token, because a fraction of a cent compounds at scale. We disagree. In the selections we advised, integration effort, commitment terms and exit cost moved the total far more than the headline rate did.
Pick on modeled year one cost plus priced exit for your real workload. The lowest rate card rarely produced the lowest total, and the pilot rate almost never survived production.
Enterprise AI procurement strategy brief
Commitment sizing, contract terms and renewal timing for AI platform deals, from the GenAI practice.
Get the white paper →How do direct model APIs, Azure OpenAI, Amazon Bedrock and Vertex AI compare on cost?
The real split is between direct model vendors and managed cloud platforms. A direct API often has the lowest unit rate and gives you more control, and you build more of the tooling yourself. A managed platform charges a margin for governance, integration and support, and saves the most effort inside the cloud you already run.
| Platform type | Unit rate | Integration effort | Best fit |
|---|---|---|---|
| Direct model API | Often lowest | Higher, you build your own | Teams wanting control |
| Azure OpenAI | Cloud rate | Lower for teams already on Azure | Microsoft heavy buyers |
| Amazon Bedrock | Cloud rate | Lower when your data and apps sit in AWS | AWS heavy buyers |
| Vertex AI | Cloud rate | Lower inside an existing Google Cloud footprint | Google Cloud buyers |
The lower integration effort and the lock in come from the same place, so price both. Google renamed Vertex AI as Gemini Enterprise Agent Platform in April 2026, and newer quotes may use that name. Our Azure OpenAI negotiation guide and Amazon Bedrock pricing guide cover each cloud's commercial terms.
How does each platform sell discounted and reserved capacity?
Each option trades a lower effective rate for slower responses or a commitment. Check current terms before you model them, because the vendors revise these offers often.
| Platform | Asynchronous or lower priority | Reserved capacity |
|---|---|---|
| OpenAI direct API | Batch API at a 50 percent discount, results within 24 hours | No published standard term; ask the account team |
| Azure OpenAI | Batch deployments; confirm the rate per model | Provisioned throughput units billed per PTU per hour, with 1 month or 1 year reservations |
| Amazon Bedrock | Batch at 50 percent below on demand for select models; Flex tier at a 50 percent discount; Priority tier at a 75 percent premium | Provisioned Throughput with no commitment, a 1 month term or a 6 month term |
| Vertex AI | Batch prediction; confirm the rate per model | Provisioned Throughput in generative AI scale units, from 1 week terms for select models to monthly and yearly terms |
Azure provisioned deployments bill for every hour they exist, whether or not tokens flow through them. Microsoft also states that reservations do not guarantee capacity. Create the deployment first to confirm capacity exists, then buy the reservation to lock in the discounted hourly rate.
Why does production inference cost so much more than the pilot?
Because a pilot runs curated requests at low concurrency, and production runs real volume, real token lengths and real concurrency. A comparison built on pilot numbers compares the wrong workload, however carefully the rate cards were normalized.
- Request volume. A pilot serves a few dozen testers. Production serves every user in scope, plus integrations that call the model with no person involved.
- Token length. Real prompts carry retrieved documents, conversation history and system instructions, all billed as input tokens on every call.
- Output length. Users ask for longer answers once they trust the tool, and output tokens cost more than input tokens on most models.
- Concurrency. Peak load decides whether you need reserved capacity at all, and a pilot never shows you the peak.
How do you check your own token volumes?
Pull real usage from the pilot before anyone quotes a production number. Every platform reports token counts.
- Azure OpenAI. Azure Monitor metrics for processed prompt tokens and generated completion tokens, per deployment.
- Amazon Bedrock. The InputTokenCount, OutputTokenCount and Invocations metrics in Amazon CloudWatch.
- Vertex AI. Cloud Monitoring for the model endpoints, with Cloud Billing reports for spend.
- Direct APIs. The usage dashboard in the vendor console, split by project or API key.
Divide tokens by requests to get the average per request, separately for input and output. Multiply by a production request count that a business owner signs off. The pilot figure is the floor for that estimate.
What does an enterprise AI platform TCO comparison look like with real numbers?
Say your pilot runs $8,000 a month in inference at list rates, and production lands at the typical four times: $32,000 a month, or $384,000 a year. Platform A is a direct API with a unit rate 15 percent lower. Platform B is a managed service inside the cloud you already run, at the standard rate.
| Line | Platform A: direct API | Platform B: managed, in your cloud |
|---|---|---|
| Inference at production volume | $326,400 | $384,000 |
| Integration and data engineering | $163,200 (50 percent, you build more) | $76,800 (20 percent, tooling in place) |
| Year one total | $489,600 | $460,800 |
| Estimated cost to exit | $60,000 | $120,000 |
| Year one total plus exit | $549,600 | $580,800 |
On year one cost alone, the platform with the higher token rate wins by $28,800. Once exit is priced, Platform A leads by $31,200. Which one is right depends on how likely you are to switch within the term, and you can only judge that with the exit number written down.
How does a spend commitment change the result?
Take the same $384,000 of usage at list and a hypothetical 20 percent discount for committed spend over 12 months. The outcome depends on which forecast you commit to.
| Commitment choice | Committed usage at list | What you pay | Against no commitment |
|---|---|---|---|
| No commitment | None | $384,000 | Baseline |
| Sized to six times pilot | $576,000 | $460,800, with $192,000 of usage unused | $76,800 more |
| Sized to three times pilot | $288,000 | $230,400 plus $96,000 at list, $326,400 in total | $57,600 less |
At a 20 percent discount, a commitment pays off only when you use at least 80 percent of it. Production running above the pilot pushes forecasts up, while abandoned use cases pull real demand down. Commit to the low end and let the upside run at standard rates.
What do integration and governance add in the first year?
Integration is the largest cost that never appears on a rate card, and it lands whichever platform wins. It covers data pipelines, access controls so the model only retrieves what each user may see, wiring into the applications people use, and logging with cost tags for chargeback.
Governance repeats every time a model version changes or a new use case goes live. Budget monitoring, evaluation and compliance review as an annual engineering line that grows with the number of use cases in production.
What does a managed platform's margin pay for?
Managed platforms carry part of the integration and governance work in their own tooling, such as content filters, retrieval services and evaluation consoles. Price that tooling against what your team would otherwise build, and against the lock in it creates.
How should you size and word an AI platform commitment?
Size the commitment to conservative demand, and write the contract so unused value is not simply lost. Forfeited capacity turns a discount into a penalty. Our guide to cloud AI commitment negotiation covers the account side.
The cheapest token often sits on the platform that costs the most to operate.
Which contract terms should you ask for?
- A ramp. Commit by quarter, rising with measured adoption, so early months are not priced on the year three forecast.
- Rollover or reallocation. Unused commitment carries forward or shifts to other models. See our swap and reallocation clause guide.
- Price hold. Rates hold for the term, including successor models. See the price hold clause language.
- A step down right. A date, such as the first anniversary, when you can reduce the commitment if a use case is abandoned. A termination for convenience clause is the fallback if the vendor refuses.
- Portable records. Prompts, evaluation sets, logs and fine tuning data returned in a usable format on exit.
- Commitment credit. Written confirmation of whether this spend counts toward any cloud commitment you already hold.
What will the account team say, and how should you answer?
| What you will hear | What to say back |
|---|---|
| Our per token rate is lower than the alternatives. | Show us the year one total at our production volume, including the integration we would build. |
| Reserve capacity now so your performance is guaranteed. | Put the guaranteed throughput in the contract. We will size the reservation after three months of production data. |
| A larger commitment earns a deeper discount. | We need to use most of it just to break even. Offer a ramp and rollover and we will discuss the size. |
| Everything you build here is portable. | List the services we would use that have no equivalent elsewhere, and price your exit assistance. |
How do you price exit cost before you sign?
Estimate the engineering work to run the same use cases on another model or platform, and put that figure beside each option before you commit. Before signing, that figure supports requests for portable records and priced exit assistance. After signing, the vendor has little reason to agree to either.
- Prompt rework. Prompts tuned for one model family rarely perform the same on another.
- Evaluation reruns. Every evaluation set must pass again before the new model goes live.
- Pipeline rebuilds. Retrieval indexes, content filters and any fine tuned models built on one cloud's services need rebuilding on the target.
- Parallel running. Both platforms bill during the cutover.
Our GenAI vendor lock in assessment walks through each item. The more managed services a use case relies on, the larger this estimate grows.
How does the answer change with your size and cloud footprint?
A company with one or two use cases and a small team usually does best on the managed platform of the cloud it already runs, because the integration saving outweighs the lock in. A large buyer with a portfolio of use cases can build its own tooling, and a direct API or second platform keeps competition alive at renewal.
Buyers split across two clouds should price both managed options and one direct API on all four drivers. Our GenAI practice runs this comparison, and the GenAI knowledge hub holds the supporting guides.
What to do next
- Define the production workload. Set request volume, token length and peak concurrency as numbers a business owner signs, and treat the pilot as a floor.
- Estimate inference at production scale. Scale measured tokens per request to production volume for each shortlisted platform at current rates.
- Add integration and governance. Budget the year one integration uplift for each platform, and governance as a recurring line.
- Size any commitment to conservative demand. Commit to the low end with a ramp and rollover, and let the upside run at standard rates.
- Price the exit. Estimate the cost of moving each use case and weigh it against the lower integration effort of staying in your current cloud.
- Negotiate with the full model. Take the four driver comparison into commercial talks so each vendor competes on your total cost.
Frequently asked questions
What actually drives enterprise AI platform cost?
Inference volume at production scale, commitment terms, integration and data engineering effort, and exit cost. The token rate is a small input to the first of these. Across the selections we advised, the rate card was rarely what decided the result once all four were modeled for the real workload instead of the pilot.
Why does the pilot understate production cost?
A pilot is built to prove that a use case works, so it runs clean test prompts for a small group at quiet times. Production adds every user in scope, longer prompts carrying retrieved content and history, and peak load. That gap is why production ran several times the pilot in the platform decisions we advised.
How much does integration add to AI platform cost?
Between 20 and 50 percent on top of platform cost in the first year, with a median of 35 percent in our file. The spread depends mostly on how much tooling the platform already provides inside your cloud and how clean your source data is before the work starts.
Are committed spend discounts on AI platforms worth taking?
Only where demand is proven. A committed discount pays off when you use enough of the commitment to cover the discount, which at 20 percent means at least 80 percent of it. Commit to the low end of the forecast and ask for a ramp and rollover.
Should we use a direct model API or a managed cloud platform?
It is a trade of control against convenience and margin. A direct API gives you more control and often a lower unit rate, and you carry more of the integration and governance build. A managed platform cuts that effort most in the cloud you already run, which is also where its lock in is highest.
Why does exit cost matter before you commit?
It measures how much lock in you are buying. None of the buyers in our file had priced it, and later switches took longer and cost more than expected. Priced before signing, it shapes the choice and the contract terms you ask for, such as portable records and exit assistance.
Does the lowest token price win at scale?
Rarely. A lower rate on the same volume saves a fixed share of the inference line, while integration effort and exit cost can each move the total by more. Compare the full year one cost plus exit for each platform, and the cheapest rate card often ends up in second place.