HomeGenAI HubAI Platform TCO
GenAI  |  Platform TCO Buyer Guide 2026

The cheapest token can sit on the most expensive platform

Token price is the smallest line in most enterprise AI budgets at scale, and it is the only one most shortlists compare. Four drivers actually decide the total: inference volume at production scale, commitment terms, integration and data engineering effort, and the exit cost nobody models until they need it. Compare platforms on all four for your real workload, and the rate card stops being the answer.

Prepared by Redress Compliance · August 10, 2026 · GenAI advisory. Based on 25 to 35 enterprise AI platform decisions, 2024 to 2025.

Executive summary

Production inference ran 3 to 6 times the pilot that justified the platform choice.

Inference volume rather than training drives the recurring cost for most buyers, and the pilot systematically understates it, because a demo runs on curated requests at low concurrency while production runs on real request volume, real token lengths, and real concurrency.

The typical scale up in our file was around four times. Any comparison built on pilot numbers is comparing the wrong workload, however carefully the rate cards were normalised.

Integration and data engineering added 20 to 50 percent on top of platform cost in the first year. The work of connecting a model to your data and systems is the real first year expense, with a median uplift of 35 percent, and it never appears on a rate card.

Governance sits alongside it: monitoring, evaluation, and compliance are recurring engineering costs rather than one time setup. Together these routinely exceed the model spend itself, which is why a platform that integrates cheaply into your existing estate can beat one with a lower unit rate.

Commitment discounts reward accurate forecasting and punish overcommitment. Committed spend lowers the unit rate and forfeits unused capacity, so the discount is only real if the demand lands.

Given that production inference runs several times the pilot in one direction and use cases get abandoned in the other, the defensible position is to size commitments to conservative demand and let the upside price at standard rates rather than buying a tier against a forecast built during a pilot.

Exit cost is real and it was never modelled, so switching later took longer and cost more than expected.

Reworking prompts, evaluations, and pipelines for a different model is genuine engineering work, and it is the price of the lock in that a managed platform inside your existing cloud estate buys you in exchange for lower integration effort. Both effects are real and both are priceable.

The choice between direct model access and a managed platform is a trade of control against convenience and margin, and it should be made with the exit number written down.

3 to 6x
Production inference cost against the pilot that justified the platform choice, typically around four times.
20 to 50%
Integration and data engineering uplift on top of platform cost in year one, a median of 35 percent.
4 drivers
Inference, commitment, integration, and exit. The token rate is one input to the first of them.
30
Enterprise AI platform decisions behind this comparison, advised across 2024 and 2025.
1.

The platforms on the cost profile that matters

Platform typeUnit rateIntegration effortBest fit
Direct model APIOften lowestHigher, build your ownTeams wanting control
Azure OpenAICloud rateLower in Azure estatesMicrosoft heavy buyers
Amazon BedrockCloud rateLower in AWS estatesAWS heavy buyers
Vertex AICloud rateLower in Google estatesGoogle Cloud buyers

The split is between direct model vendors and managed cloud platforms, and each side trades the same two things in opposite directions.

Direct API access gives control and often the lowest unit rate with less managed tooling around it, so the buyer carries more of the integration and governance build.

Managed platforms add governance, integration, and support at a margin, and they lower integration effort most inside the cloud estate you already run, which is also where they raise lock in most. Price both effects rather than treating one as a feature and the other as a footnote.

The Azure commercial detail sits in the Azure OpenAI negotiation guide and the AWS side in the Bedrock pricing guide.

2.

The costs that are not on any rate card

Free white paper

The enterprise AI procurement strategy brief

Sourcing, contracting, and renewal across the AI platform layer, with the commitment arithmetic and the data terms that decide whether a deal is safe to sign.

Get the white paper →
3.

Choosing on total cost rather than rate card

The method is unglamorous and it is the whole difference between a defensible decision and a plausible one. Start by defining the real production workload rather than the pilot: request volume, token length per request, and concurrency at peak, stated as numbers somebody owns rather than as a range.

Estimate inference cost at that volume, which in our file ran three to six times what the pilot implied, and treat the pilot figure as a floor rather than an estimate.

Add integration and data engineering effort to the first year total, because that work commonly adds 20 to 50 percent on top of platform spend and it lands whichever platform wins.

Size any commitment to conservative, defensible demand, since a committed rate is only a discount if the volume arrives and forfeited capacity converts the discount into a penalty.

Then model the exit cost of moving to another platform, which is the number that decides how much lock in you are actually buying when you choose the platform that sits inside your existing cloud estate.

Only with all four modelled does the comparison mean anything, and the ranking frequently inverts at that point, because the cheapest token often sits on the platform that costs most to operate.

Weigh the lock in of staying inside your existing estate explicitly rather than accepting it as convenience, and price the governance layer as recurring rather than as setup. The full library sits in the GenAI knowledge hub.

Try Vera AI · free 30 day trial
Vera models production inference against your real workload, prices integration and exit alongside the rate card, and benchmarks the commitment before you sign it.
  • Percentile standing for your exact deal size and industry, from real closed transactions
  • Scenario simulation before the call: test alternative terms and see the financial impact of each
  • A negotiation playbook, talking points, and a two page executive brief on day one
Start the free Vera AI trial →30 days free · no credit card · cancel anytime
4.

What we saw across AI platform selections, 2024 to 2025

The common advice is to pick the platform with the lowest token price, because at scale a fraction of a cent per token compounds into the biggest number. We disagree, because in the selections we advised the token rate was rarely the deciding factor once the full picture was modelled:

4x
Pilot to production

Typical scale up in inference cost between the pilot that justified the platform and the production workload that followed it.

35%
First year integration uplift

Median addition to platform cost from integration and data engineering work in year one, on top of every rate card comparison.

Three patterns recurred: inference at production scale running 3 to 6 times the cost of the pilot that justified the platform choice, integration and data engineering adding 20 to 50 percent on top of platform cost in the first year, and exit cost never being modelled at all.

So a later switch took far longer and cost more than buyers expected.

The buyer side move is to model total cost across inference, commitment, integration, and exit for your actual workload, then choose. The lowest rate card rarely produces the lowest total, and the pilot rate almost never survives production.

5.

Your first five moves

  1. Define the real production workload, request volume, token length, and concurrency at peak, and treat the pilot number as a floor rather than an estimate.
  2. Estimate inference at production scale, not pilot scale, because it ran 3 to 6 times higher across the selections we advised and it dominates the recurring bill.
  3. Add integration and data engineering to the first year total, a median 35 percent uplift, since that work lands whichever platform you choose.
  4. Size any commitment to conservative demand, letting the upside price at standard rates, because forfeited capacity turns a discount into a penalty.
  5. Model the exit cost before committing, then weigh it against the lower integration effort of staying inside your existing cloud estate. The GenAI practice builds the model with you.
6.

Frequently asked questions

What actually drives enterprise AI platform cost?

Four things: inference volume at production scale, commitment terms, integration and data engineering effort, and exit cost. The token rate is a small input to the first of them.

In the selections we advised, the rate card was rarely the deciding factor once all four were modelled for the buyer's real workload rather than for the pilot.

Why does the pilot understate production cost?

Because a pilot runs curated requests at low concurrency while production runs real volume, real token lengths, and real concurrency. Across our file, production inference ran 3 to 6 times the pilot that justified the platform choice, typically around four times.

A comparison built on pilot numbers is comparing the wrong workload however carefully the rates were normalised.

How much does integration add?

Between 20 and 50 percent on top of platform cost in the first year, a median of 35 percent. Connecting the model to your data and systems is the real first year expense, and it never appears on a rate card.

Governance work, meaning monitoring, evaluation, and compliance, sits alongside it as a recurring engineering line rather than a setup task.

Are committed spend discounts worth taking?

Only where the demand is defensible. Commitment lowers the unit rate and forfeits unused capacity, so the discount is real only if the volume lands.

Given that production inference runs several times the pilot in one direction and use cases get abandoned in the other, size commitments to conservative demand and let the upside price at standard rates.

Should we use a direct model API or a managed cloud platform?

It is a trade of control against convenience and margin. Direct API access gives control and often the lowest unit rate with less tooling, so you carry more of the integration and governance build.

Managed platforms lower integration effort most inside the cloud estate you already run, which is also where they raise lock in most. Price both effects.

Why does exit cost matter before you commit?

Because reworking prompts, evaluations, and pipelines for a different model is genuine engineering work, and it is the measure of how much lock in you are buying. It was never modelled in our file, so later switches took longer and cost more than expected.

Modelled before commitment it is a number; modelled after, it is leverage you no longer have.

Does the lowest token price win at scale?

Rarely. The cheapest token often sits on the platform that costs most to operate, because integration effort, commitment terms, and exit cost moved the total far more than the headline rate in the selections we advised.

Compare on total cost across all four drivers for your actual workload, and the ranking frequently inverts.

© 2026 Redress Compliance · Independent, buyer sideredresscompliance.com
Industry Recognized
500+ Enterprise Clients
$2B+ Under Advisory
11 Vendor Practices
100% Buyer Side Independent
GenAI White Paper

The full enterprise AI procurement strategy brief from the GenAI practice.

Sourcing, contracting, and renewal across the AI platform layer, with the commitment arithmetic and the data terms that decide whether a deal is safe to sign.

Gated with a work email on the download page. No sales follow up you did not ask for.

Get the White Paper →
Independent, buyer side. We never share your details with vendors.
Run the software spend health check against your AI estate in under five minutes.
Open the Tool → GenAI Practice →
Editorial boardroom interior

The advisor your vendors do not want.

500+ enterprise clients. 11 vendor practices. Industry recognized. One conversation can change what you pay for the next three years.

Stay ahead of AI platform pricing and contract moves.

One buyer side briefing a week. Renewal signals, discount bands, and the levers that work. No vendor spin.