The token is the meter, and the outcome is the only benchmark
List price per token is published; realized cost per useful output, after gateway markup, model routing, retries, and fine tuning, is not: at enterprise volume on frontier tier models, realized blended cost lands between $3 and $12 per million tokens, and the list rate is the floor of that band rather than the middle. Two buyers on the same vendor and the same rate card can end two or three times apart on cost per outcome, because routing, prompt size, retries, and tier choice do more work than the headline rate ever does.
Prepared by Redress Compliance · August 8, 2026 · GenAI advisory. Based on 25 to 35 enterprise AI deployments and gateway reviews supported 2024 to 2025.
Executive summary
The gateway layer is a contract wearing a tool's clothes, adding 15 to 35 percent.
Most enterprises do not call the model APIs directly in production, and gateway and reseller layers added 15 to 35 percent above raw API list on most contracts we read, with managed wrappers running higher: the layer buys routing, observability, and governance.
And it prices like a reseller margin, which makes it a commercial term to negotiate rather than an infrastructure choice to default.
Treat the gateway as a contract, benchmark its markup, and price the thin pass through against the managed wrapper deliberately.
The realized bands run three tiers, and the estate spans all of them.
At production volume the realized blended cost runs roughly $0.50 to $2 per million tokens on small models for chat, $2 to $5 on midweight reasoning for most enterprise tasks, and $5 to $12 on frontier tier for agentic and complex work: a mixed estate weights toward the most expensive layer.
And the drivers of where a buyer lands inside each band are token count discipline, model selection per task, retry policy, and the gateway margin, never the rate card everyone compared on.
Retry waste consumed 10 to 25 percent of agent volume before any value. Retry, reroute, and tool call waste consumed 10 to 25 percent of agent token volume before any user value was produced.
And fine tuning and routing changed bills more than list rates did: the cheapest list price often ships the most expensive bill, with cost per useful output varying 2 to 4 times across stacks that had picked the cheapest published rate, because routing and prompt structure swamped the headline.
Chat is the cheapest workload, agents the most expensive, batch in the middle.
Cost per useful output is the only benchmark a CFO can act on.
The vendor lists have converged in shape and parted in detail, OpenAI, Anthropic, Google Vertex, Azure OpenAI, and AWS Bedrock differing more in discount structure and term than headline rate, with tier mix, context windows.
And output rates mattering more than input pricing: the buyer side program benchmarks per outcome at production volume, negotiates the commit and routing terms that compound.
And runs the routing audit that forces the math into the open, because buyers who treated each request as a costed transaction landed in the lower half of every band, and buyers who treated tokens as free until the invoice arrived landed at double their plan.
The realized bands, list against loaded
| Tier | Realized blended band | The workload it carries |
|---|---|---|
| Small and mini models | $0.50 to $2 per million tokens | Chat, retrieval, and cheap classification |
| Midweight reasoning | $2 to $5 per million | The workhorse of most enterprise tasks |
| Frontier tier | $5 to $12 per million | Agentic and complex reasoning workloads |
| The agentic stack, loaded | The top of every band plus the waste | Where retries and tool calls multiply the meter |
The list is the input, not the answer.
The realized figure loads the gateway markup, the retry overhead.
And the routing choices onto the published rate, which is why the raw API comparison that procurement ran predicts almost nothing about the bill: a mixed estate using small models for retrieval and frontier for synthesis spans all three bands and weights toward the expensive layer.
And the two buyers on identical rate cards who land two to three times apart are separated entirely by discipline, not by negotiation.
The buyer program, outcome benchmarked
- Benchmark per outcome at production volume: the cost of a resolved ticket, a completed draft, or a closed task, the unit a CFO can act on, never the token.
- Audit the routing: the written policy of which model serves which task, the discipline worth 20 to 35 percent in the routing files.
- Meter the retries: the 10 to 25 percent of agent volume producing nothing, instrumented and capped before it compounds.
- Negotiate the gateway as a contract: the 15 to 35 percent markup benchmarked, the thin pass through priced against the managed wrapper.
- Match commits to measured baselines: the vendor tiers differing in structure and term more than rate, and the commitment earning its discount only against real volume.
The enterprise AI contract negotiation playbook
The gateway terms, the commit structures, and the routing clauses that decide the realized rate, worked across the model vendors.
Get the white paper →The vendor shapes, converged and parted
The five channels price the same idea differently: OpenAI's published ladder spans cheap small models to frontier output in the tens of dollars per million, Anthropic's card mirrors the shape with the fast against frontier gap.
Vertex prices per character and token across Gemini tiers with separate grounding and tool call lines, Bedrock publishes per model rates plus the provisioned throughput line that changes the math at scale.
And Azure OpenAI wraps the same models in enterprise agreement terms that often replace the public rate entirely.
The structural read is that tier mix, context window, and output rate now matter more than the headline input rate everywhere.
The per vendor depth runs in the Anthropic pricing history, where the output token drove 60 to 80 percent of bills, the consumption governance in the token cost surge report, the attach side of the same estate in the GenAI pricing report, and the off books demand in the shadow AI report.
- Percentile standing for your exact deal size and industry, from real closed transactions
- Scenario simulation before the call: test alternative terms and see the financial impact of each
- A negotiation playbook, talking points, and a two page executive brief on day one
What we saw across AI workloads, 2024 to 2025
Across roughly 25 to 35 enterprise AI deployments and gateway reviews our team supported between 2024 and 2025, the realized cost per useful output was rarely close to the API list price the buyer had compared on:
Above raw API list on most contracts read, higher on managed wrappers.
Cost per useful output across stacks that picked the cheapest published rate.
The report's closing distinction is behavioral: the buyers who treated each request as a costed transaction, with the routing policy written, the retries metered, and the gateway margin negotiated, landed in the lower half of every band.
And the buyers who treated tokens as free until the invoice arrived landed at double their plan, rarely connecting the two until a routing audit forced the math into the open.
The workload hierarchy holds everywhere, chat cheapest, batch in the middle, agents most expensive, and the benchmark that survives every vendor list change is the same one that started the report: the outcome, not the token.
Your first five moves
- Benchmark cost per useful output at production volume, the only unit the CFO can act on.
- Run the routing audit and write the policy down, where 2 to 4 times of outcome cost hid.
- Meter and cap the retry waste, the 10 to 25 percent of agent volume producing nothing.
- Negotiate the gateway markup as a contract term, the 15 to 35 percent riding above list.
- Commit against measured volume, structured by vendor, since the tiers differ in term more than rate. The cost optimization practice runs the audit with you.
Frequently asked questions
What do enterprises actually pay per million AI tokens?
Between $3 and $12 per million blended tokens at enterprise volume on frontier tier models, with small models at $0.50 to $2 and midweight reasoning at $2 to $5: the published API list is the floor of the band, not the middle, because gateway markup, retries, and routing overhead load onto it.
Two buyers on the same rate card routinely land two to three times apart on cost per outcome.
How much do AI gateways add to costs?
15 to 35 percent above raw API list on most enterprise contracts we benchmarked, with managed wrappers running higher: the layer delivers routing, observability, and governance, and it prices like a reseller margin, which makes it a commercial term to negotiate rather than an infrastructure default.
The thin pass through against the managed wrapper is a priced decision, not an architectural one.
What is AI retry waste?
The token volume consumed by retries, reroutes, and failed tool calls before any user value is produced: 10 to 25 percent of agent token volume in our reviews, and one of the two levers, alongside routing, that change bills more than any list rate.
It is instrumented and capped as an engineering discipline, and invisible until a routing audit forces the math into the open.
Why is the cheapest AI model often the most expensive?
Because cost per useful output, not cost per token, is the real bill: stacks that picked the cheapest published rate varied 2 to 4 times on outcome cost, as undersized models generated retries, longer prompts, and lower first pass quality that swamped the headline saving.
The benchmark unit is the resolved ticket or the completed task at production volume, the number a CFO can actually act on.
How do the AI model vendors' prices compare?
They converge in shape and part in detail: OpenAI, Anthropic, Google Vertex, Azure OpenAI, and AWS Bedrock differ more in discount structure, commit terms, and channel mechanics than in headline rate, with tier mix, context windows, and output rates mattering more than input pricing everywhere.
The comparison worth running is each vendor's routed configuration against your task mix at measured volume, never the rate cards side by side.
Which AI workloads cost the most?
Agents, by a wide margin: chat is the cheapest workload, batch processing sits in the middle, and agentic workloads run most expensive because they consume frontier tier output heavily and carry the 10 to 25 percent retry and tool call waste on top.
The estate that routes by workload, small models for retrieval, midweight for standard tasks, frontier only where reasoning demands it, holds the blend down across all three.