Bedrock spend sat outside the EDP commitment math, leaving 10 to 25 percent of leverage unused
Bedrock prices generative AI on three commercial shapes and none of them behave like the rest of an AWS bill. The fastest growing and least governed line in the invoice is also the one most often left out of the commitment that prices everything else.
Prepared by Redress Compliance · August 16, 2026 · AWS advisory. 20 to 30 Bedrock and EDP reviews, 2024 to 2025.
Executive summary
Bedrock counts toward the EDP commitment and was routinely left out of the math, costing 10 to 25 percent of negotiating leverage in the reviews we ran. The spend was real; it simply was not in the spreadsheet that set the commitment.
Model choice drove a 5 to 15 times cost spread for near identical task quality, and it was unaudited in almost every estate. The catalogue spread on a blended workload is wider still.
On demand ran 30 to 60 percent higher than provisioned throughput for steady workloads that never moved to commitment, which is the opposite of the usual advice that most workloads should stay on demand.
Output tokens cost three to five times input tokens, so answer length is a pricing decision. A verbose system prompt is cheap and a verbose model is not.
Three commercial shapes, and what each one actually charges for
Bedrock hosts foundation models from Anthropic, Meta, Mistral, Cohere, Amazon, AI21, and Stability under a single AWS contract, and prices inference three different ways. Each carries different cost dynamics, and mixing them up is where budgets fail.
| Model tier | Input per million tokens | Output per million tokens | Typical use case |
|---|---|---|---|
| Premium reasoning | $3.00 to $15.00 | $15.00 to $75.00 | Long context analysis, agent workflows |
| Standard chat | $0.80 to $3.00 | $4.00 to $15.00 | General chat, summarisation |
| Lightweight | $0.20 to $0.80 | $1.00 to $4.00 | Classification, retrieval, routing |
| Open weight | $0.30 to $1.50 | $0.60 to $3.00 | Self hosted or replicated workloads |
| Embedding | $0.02 to $0.20 | Not applicable | Retrieval augmented generation |
Run the arithmetic on a blended workload rather than on a headline rate. Take a task with roughly equal input and output tokens. On premium reasoning at the top of the band that is $15 plus $75, or $90 per million token pairs. On lightweight it is $0.80 plus $4.00, or $4.80. That is close to nineteen times for the same volume of work, before anyone has assessed whether the premium model produced a better answer. The audited figure, comparing models that genuinely delivered the same task quality, was 5 to 15 times, which is the number to plan against.
The commitment math, and the line that gets left out
- Bedrock spend counts toward the AWS Enterprise Discount Program commitment, and excluding it from the commitment math left 10 to 25 percent of leverage unused in the reviews we ran.
- EDP discounts step by commitment threshold and do not exceed 20 percent: roughly 5 to 10 percent between $1m and $5m, 10 to 15 percent between $5m and $25m, and 15 to 20 percent above $25m. Moving a threshold is worth more than arguing inside a band.
- Growing Bedrock is therefore a commitment argument, not just an AI budget question. A forecast that includes it can cross a threshold the same forecast without it does not reach.
- The marketplace contribution cap toward commit is a separate mechanism, not a discount, and should not be blended into the same number when you model the ladder.
- Offers mix credits and discounts and change quarter to quarter, so price the structure across the term rather than comparing this quarter's headline against last quarter's.
- Pin the model version in the contract. AWS deprecates older versions on a rolling schedule, and a forced migration mid term is a cost and a quality risk you did not price.
The AWS EDP negotiation brief
The commitment ladder, the discount bands by threshold, and the buyer side moves that decide an Enterprise Discount Program renewal.
Get the brief →The governance gap is upstream of the pricing question
Most Bedrock cost advice starts with the rate card and works down: pick the cheap model, watch the output tokens, consider provisioned throughput. That is all correct and it is not where the money was lost in the estates we reviewed. Bedrock token spend was the fastest growing and least governed line in the AWS bill, and the governance gap sat upstream of every pricing decision, in the question of whether the spend was visible to the people negotiating the contract at all.
The mechanism is ordinary and expensive. Bedrock arrives as an engineering decision rather than a procurement one, billed inside an existing AWS account, at a scale that is trivial during a pilot and material within two quarters. The commitment forecast that sets an Enterprise Discount Program term is usually built from the historical bill and the known roadmap, and a line that was rounding error at forecast time and is now a growth curve does not appear in either. So the commitment gets set on a smaller number than the estate will actually spend, and the buyer discovers later that the spend which would have moved them up a threshold was already happening.
The second gap is the model audit that nobody runs. Across the reviews, model choice drove a 5 to 15 times cost spread for near identical task quality, and almost nobody had tested it. The reason is structural rather than careless: the team that selects a model does so during a build, optimises for output quality under time pressure, and has no visibility of the marginal cost of a token. Once the workload is in production the selection is rarely revisited, because revisiting it means re running evaluations against a system that currently works. Running three to five models over the same prompts and comparing quality, latency, and cost is a day of work that recurs annually and pays for itself in most estates on the first pass.
The provisioned throughput question then splits in a way worth stating carefully, because the honest answer contains a contradiction. Most enterprise Bedrock workloads burst and sit idle, so on demand genuinely wins for them, and the breakeven sits around forty percent utilisation of reserved capacity. But steady workloads that never moved to commitment paid 30 to 60 percent more than provisioned throughput would have cost. Both are true. The discipline is to segment by utilisation profile rather than adopting a single posture: measure sustained throughput per workload, commit only the portion that clears forty percent utilisation, and leave the bursty remainder on demand. The commitment ladder itself is worked in the EDP discount benchmarks, and the wider library sits in the AWS practice.
- Your quote benchmarked against 500,000+ real closed deals, adjusted for size, region, and industry
- Commitment thresholds modelled so you see which one the corrected forecast reaches
- Every risky clause flagged with the exact quote, the page, and the replacement language
Provisioned throughput, and when it genuinely wins
- PTU reserves model capacity at a fixed monthly fee, billed per model unit per hour, on one month or six month commitment terms.
- The six month commitment runs roughly forty percent below the one month rate, so term length is the largest single lever inside the PTU decision.
- Breakeven sits around forty percent utilisation of the reserved capacity for most models. Below that, on demand wins and PTU is a stranded cost.
- Most enterprise workloads should not be on PTU. They burst on demand and sit idle most hours, and PTU is a tool for sustained production inference at high volume rather than a default posture.
- Steady workloads that stayed on demand paid 30 to 60 percent more than provisioned throughput would have cost, which is the mirror image of the same discipline.
- Fine tuning carries a one off training charge plus a higher per token inference rate on the customised model, so model customisation has to clear both costs rather than only the training line.
What the Bedrock reviews showed, 2024 to 2025
Across roughly 20 to 30 AWS Bedrock and EDP reviews, Bedrock token spend was the fastest growing and least governed line in the bill:
Cost difference across models delivering near identical task quality, unaudited in almost every estate reviewed.
Share of negotiating leverage lost by excluding Bedrock spend from the EDP commitment math.
The third pattern was the on demand premium. Steady workloads that never moved to commitment ran 30 to 60 percent higher than provisioned throughput would have cost, while the majority of bursty workloads were correctly left on demand. Segmenting by utilisation profile is what separates those two populations.
Two mechanical traps sit underneath all of it. Output tokens cost three to five times input tokens, so answer length is a pricing decision rather than a formatting one. And Bedrock pricing varies by region with some models available only in a single region, which turns a residency requirement into a rate difference.
Watch the briefing · 4:10Negotiating AWS 3: The Discount StackHow commitment thresholds, Savings Plans, and the negotiated rate multiply rather than add.
Your first five moves
- Add Bedrock to the EDP commitment forecast with its actual growth curve, then check whether the corrected total crosses a threshold the old forecast did not reach.
- Run three to five models over the same prompts and compare quality, latency, and cost, then make that evaluation an annual exercise rather than a one time build decision.
- Segment workloads by sustained utilisation and commit only the portion clearing roughly forty percent of reserved capacity, leaving the bursty remainder on demand.
- Constrain output length where quality allows, because output tokens cost three to five times input and answer verbosity is the cheapest lever to pull.
- Pin model versions and check regional availability in the contract, so a rolling deprecation does not force an unpriced migration. The AWS practice rebuilds the commitment model with you.
Frequently asked questions
How does AWS Bedrock pricing work?
Three commercial shapes. On demand charges per million tokens consumed, with separate input and output rates per model. Provisioned Throughput Units reserve capacity at a fixed monthly fee on one or six month terms. Model customisation carries a one off training charge plus a higher inference rate on the fine tuned model.
Does Bedrock spend count toward an AWS EDP commitment?
Yes, and that is the point most often missed. Across the reviews we ran, Bedrock was routinely excluded from the commitment math even though it counts, which left 10 to 25 percent of negotiating leverage unused because the forecast understated total spend.
How much does model choice actually matter?
A great deal. Model choice drove a 5 to 15 times cost spread for near identical task quality, and it was unaudited in almost every estate. On raw rate card a blended workload can differ by close to nineteen times between premium reasoning and lightweight models.
Why do output tokens cost more than input?
They are priced separately and run three to five times the input rate on the same model. That makes answer length a pricing decision: a long system prompt is comparatively cheap, and a model that answers at length is not. Constraining output where quality allows is the cheapest available lever.
Should we move to Provisioned Throughput Units?
Usually not. Most enterprise Bedrock workloads burst and sit idle most hours, and breakeven sits around forty percent utilisation of the reserved capacity. PTU is a tool for sustained production inference at high volume, not a default posture.
Then why did steady workloads overpay on demand?
Because the discipline is per workload rather than per estate. Steady workloads that never moved to commitment ran 30 to 60 percent above what provisioned throughput would have cost. Segment by sustained utilisation and commit only the portion that clears the breakeven.
How much does the PTU term length matter?
The six month commitment runs roughly forty percent below the one month rate, which makes term the largest single lever inside the PTU decision once you have established that PTU is right for that workload at all.
What EDP discount should we expect?
Discounts step by commitment threshold and do not exceed 20 percent: roughly 5 to 10 percent between $1m and $5m, 10 to 15 percent between $5m and $25m, and 15 to 20 percent above $25m. Moving a threshold is worth more than arguing within a band, which is why the Bedrock forecast matters.
Is the marketplace contribution a discount?
No. The cap on marketplace spend contributing toward commitment is a separate mechanism from the discount ladder, and blending the two produces a misleading model. Treat it as a way to reach a threshold, not as a rate reduction.
What should we pin in the contract?
The model version and the regional availability. AWS deprecates older model versions on a rolling schedule, and Bedrock pricing varies by region with some models available only in one. Both turn into unpriced migrations if they are left to the roadmap rather than written down.
Negotiating AWS 5: Know What Good Looks Like
Published benchmark tables disagree by a factor of three. The one verifiable AWS discount is 9 percent, filed with regulators. Plus the bands, the breakpoints, competition worth 3 to 8 points, and effective against headline rate.