A data center aisle lined with server racks
AWS Bedrock

AWS Bedrock pricing in 2026. What tokens, throughput and the EDP cost you.

Token rates by model tier, worked monthly costs, when Provisioned Throughput pays, and how Bedrock spend changes the AWS Enterprise Discount Program commitment.

Contact Us AWS Advisory
500+Enterprise clients
$2B+Under advisory
PublishedJuly 18, 2025UpdatedSeptember 24, 2026
ContentsKey takeawaysHow Bedrock pricing worksToken rates by model tierModel choice and costBedrock and the EDPProvisioned ThroughputWhat we have seenNegotiating with AWSWhat to do nextFAQ

Bedrock bills per token, per reserved model unit and per customization, and none of it behaves like the rest of an AWS bill. The costliest gaps sit in model choice and in the EDP commitment forecast.

Key takeaways
  • Three commercial shapes. On demand tokens, Provisioned Throughput by the hour, and customization with its own training charge each need a separate cost decision.
  • Model choice outweighs discounts. Models of near identical quality differed in cost by a wide multiple in our reviews, and almost no team had tested the cheaper options.
  • Output length is a price. On most models output tokens cost several times input, so capping answer length is the quickest saving available.
  • Bedrock belongs in the EDP forecast. It counts toward the commitment, and adding its growth curve can move you across a discount threshold.
  • Commit per workload. Reserve throughput only for workloads busy enough to clear the breakeven, and leave bursty traffic on demand.
  • Pin versions and regions. Model deprecation, extended access pricing and single region models all turn into unpriced migrations unless the contract covers them.

Amazon Bedrock charges for generative AI in ways that behave nothing like the rest of an AWS bill. You pay per million tokens, and on most models output is priced well above input. You can reserve model capacity by the hour. Customizing a model adds a training charge and changes how the result can be served.

In our Bedrock and Enterprise Discount Program reviews, the largest losses came from decisions made far from procurement: which model a team picked, whether Bedrock was counted in the EDP commitment, and which workloads stayed on demand. This guide covers each, with worked numbers you can rerun on your own bill.

How does Amazon Bedrock pricing work in 2026?

Bedrock prices inference three ways. On demand charges per million tokens consumed, with separate input and output rates for each model. Provisioned Throughput reserves model capacity at a fixed hourly price. Model customization carries a one off training charge, then runs inference on the customized model at a higher cost than the base model on demand.

Around those three shapes, AWS has added options that change the effective rate without a new contract. Each one is a setting your engineers choose per request or per workload, which is why finance rarely sees them.

Bedrock pricing options and when each one fits
OptionHow it billsCommitmentFits
On demand, Standard tierPer million input and output tokensNonePilots, bursty and unpredictable traffic
Priority tierPer token at a 75 percent premium to Standard on the models that offer itNoneCustomer facing calls where latency matters more than cost
Flex tierPer token at a 50 percent discount to Standard on the models that offer itNoneEvaluations, summarization and jobs that tolerate slower responses
Batch inferencePer token, 50 percent lower than on demand for select modelsNoneOvernight document runs, classification backlogs
Provisioned ThroughputHourly per model unitNone, 1 month or 6 monthsSustained, high volume production inference and most custom models
Reserved tierFixed price per 1K tokens per minute, billed monthly1 month or 3 months, through your account teamWorkloads that need guaranteed capacity; unused reservation is still billed, and overflow runs at Standard rates

Prompt caching sits on top of these. AWS says it can cut cost by up to 90 percent on repeated prompt content, such as a long system prompt or a reference document sent with every call. It is the cheapest saving on this list, and it needs no commercial negotiation at all.

Which models does Bedrock host, and who bills you for them?

Bedrock hosts foundation models from Anthropic, Meta, Mistral, Cohere, Amazon, AI21, and Stability, among others, all reached through your AWS account. The billing path is less uniform. AWS states that some serverless models are sold by third party providers and appear on an AWS Marketplace bill, and its service terms list Anthropic among those providers.

That split matters for the EDP commitment. Pull a month of billing data and check which part of your Bedrock spend arrives as Amazon Bedrock and which part as Marketplace charges.

Watch the briefingEpisode 5 of 12 · 4:40

What do Bedrock models cost per million tokens?

List rates run from a few cents to $75 per million tokens, and output usually costs more than input. The bands below group the current catalogue by what enterprises use each tier for. Rates vary by model and by region, so treat them as planning ranges and confirm the exact model on the Bedrock pricing page.

Bedrock on demand rates by model tier
Model tierInput per million tokensOutput per million tokensTypical use case
Premium reasoning$3.00 to $15.00$15.00 to $75.00Long context analysis, agent workflows
Standard chat$0.80 to $3.00$4.00 to $15.00General chat, summarization
Lightweight$0.20 to $0.80$1.00 to $4.00Classification, retrieval, routing
Open weight$0.30 to $1.50$0.60 to $3.00Self hosted or replicated workloads
Embedding$0.02 to $0.20Not applicableRetrieval augmented generation

Why do output tokens cost more than input?

Providers price generation above reading. On most commercial models the output rate runs three to five times the input rate, although some open weight models, such as Llama 3.3 70B, charge the same for both.

That makes answer length a pricing decision rather than a formatting one. A verbose system prompt is cheap, especially once it is cached, while a model that answers at length costs more on every call.

What does a real workload cost on each tier?

Take a task with roughly equal input and output tokens. At the top of the premium reasoning band that is $15 plus $75, or $90 per million token pairs. On lightweight it is $0.80 plus $4.00, or $4.80. That gap of close to nineteen times covers the same volume of work, before anyone checks whether the premium answer was better.

Now take a hypothetical support assistant handling 2 million requests a month, each with 1,500 input tokens and 400 output tokens. That is 3,000 million input tokens and 800 million output tokens a month. The table prices it at the top of three bands.

Worked example: 2 million requests a month (hypothetical)
Tier, top of bandInput costOutput costMonthly total
Premium reasoning ($15.00 / $75.00)$45,000$60,000$105,000
Standard chat ($3.00 / $15.00)$9,000$12,000$21,000
Lightweight ($0.80 / $4.00)$2,400$3,200$5,600
Standard chat, answers capped at 250 tokens$9,000$7,500$16,500

On standard chat, output is about a fifth of the tokens and 57 percent of the bill. Cutting average answers from 400 to 250 tokens saves $4,500 a month, or 21 percent, with no model change or contract work. Where the model supports batch and the traffic can wait overnight, batch pricing halves either figure.

Does the rate card stay still between evaluations?

It changes in both directions. Anthropic cut Opus list pricing from $15 and $75 per million tokens to $5 and $25 with Opus 4.5, so a model rejected last year on cost may now be cheaper. Older versions can go up: Bedrock lists Claude 3.5 Sonnet on public extended access at $6.00 input and $30.00 output, twice its original rate.

Every Bedrock model passes through Active, Legacy and End of Life states. Most models get a 6 month Legacy period and some only 45 days, and AWS says the provider may adjust pricing during extended access. A workload pinned to an old version therefore carries a price risk as well as a migration deadline.

Free white paper

AWS EDP negotiation guide

Size the next commitment with Bedrock in it: sizing worksheet, ramps, Marketplace wording and benchmark bands.

Get the white paper →

How much does model choice change the Bedrock bill?

It changes the bill more than any discount you will negotiate. In our reviews, model choice drove a 5 to 15 times cost spread between models that delivered near identical task quality, and in almost every account that spread had never been tested. The raw rate card gap is wider, but plan against the tested spread.

The cause is structural. The team that selects a model does it during a build, optimizes for output quality under time pressure, and cannot see the marginal cost of a token. Once the workload is in production the choice is rarely revisited, because that means rerunning evaluations against a system that works.

How do you run a model cost review?

  1. Sample real traffic. Export a set of production prompts per workload, including the long and awkward ones, with the answers users accepted.
  2. Pick three to five candidates. Include the current model, one cheaper model from the same provider, and at least one model from another provider.
  3. Score quality, latency and cost together. Amazon Bedrock Evaluations can run automated and human judged comparisons. Price each run with the actual input and output token counts it produced.
  4. Route by task. Send classification and routing to lightweight models and keep the premium model for the steps that need it. Bedrock Intelligent Prompt Routing can do this within a model family.
  5. Repeat every year. The review is about a day of work, and it paid for itself on the first pass in most accounts we reviewed.

Does Bedrock spend count toward an AWS EDP commitment?

Yes. Bedrock spend counts toward the AWS Enterprise Discount Program commitment, yet it was routinely left out of the commitment forecast. In the reviews we ran, that omission left 10 to 25 percent of the buyer's negotiating room unused, because the forecast understated what the company was about to spend.

The mechanism is ordinary. Bedrock arrives as an engineering decision, billed inside an existing AWS account, at a scale that is trivial during a pilot and material within two quarters. The EDP forecast is built from the historical bill and the known roadmap, and a line that was rounding error at forecast time appears in neither.

Observed EDP discount bands by annual commitment
Annual commitmentTypical discount
$1m to $5m5 to 10 percent
$5m to $25m10 to 15 percent
Above $25m15 to 20 percent

Discounts step by commitment threshold and do not exceed 20 percent. Crossing a threshold is worth more than arguing inside a band, so a growing Bedrock line is a commitment question as much as an AI budget one. Our EDP discount benchmarks cover the bands in more detail.

Worked example: when Bedrock pushes you across a threshold

Say your finance team forecasts $4.3m of annual AWS spend for the next term, built from last year's bill. Bedrock is not in it, yet it now runs at $75,000 a month, or $900,000 a year. The corrected forecast is $5.2m, which takes the negotiation from the lowest band into the second.

On $5.2m of spend, each extra discount point is worth $52,000 a year. Commit only the Bedrock spend you would still incur if one project stalled, and read our guidance on EDP shortfall risk before you raise the number.

Is Marketplace billed model usage part of the same calculation?

Treat it separately. The cap on how much Marketplace spend contributes toward the commitment is its own mechanism, not a discount, and blending it into the discount ladder produces a misleading model. Because some Bedrock models bill through Marketplace, check how your qualifying spend definition treats them. Our note on Marketplace spend in the EDP covers the drafting.

AWS offers also mix credits and discounts, and the mix changes quarter to quarter. Price the whole structure across the term, and compare offers on that total.

When is Bedrock Provisioned Throughput cheaper than on demand?

It is cheaper when a workload keeps the reserved capacity busy. Provisioned Throughput is billed per model unit per hour on a 1 month or 6 month commitment, or with no commitment at a higher hourly rate. For most models, breakeven sits around forty percent use of the reserved capacity, and below it the reservation is stranded cost.

  • Term length. The six month commitment runs roughly forty percent below the one month rate, the largest single price difference inside the throughput decision.
  • Bursty traffic. Most enterprise Bedrock workloads burst and sit idle most hours, so on demand is the right answer for them.
  • Steady traffic. Steady workloads that stayed on demand paid 30 to 60 percent more than Provisioned Throughput would have cost.
  • Custom models. AWS requires Provisioned Throughput to serve most customized models. Custom Amazon Nova and Meta Llama models can also be deployed on demand.

Should every Bedrock workload stay on demand?

The usual advice is to keep Bedrock on demand and avoid commitments until usage settles. We disagree with it as a blanket rule, because the steady workloads in our reviews paid a premium for that caution. Measure sustained throughput per workload, commit only the portion that clears the breakeven, and leave the bursty remainder on demand.

What does fine tuning add to the cost?

Customization carries a one off training charge plus monthly storage for the custom model. Inference then usually runs on Provisioned Throughput, which for a low volume workload costs far more per token than the base model on demand. A fine tuned model has to beat both costs combined, measured against a cheaper base model with a better prompt.

How do you check your own usage profile?

  • CloudWatch. The AWS/Bedrock namespace reports Invocations, InputTokenCount and OutputTokenCount per model. Plot tokens per minute across a month to see whether a workload is steady or bursty.
  • Cost Explorer. Filter on the Amazon Bedrock service and group by usage type to separate input tokens, output tokens and Provisioned Throughput hours.
  • Application inference profiles. Tag a profile per application so each team's token spend shows up in cost allocation reports.
  • Your account team. Ask for the tokens per minute each model unit delivers for your model, since AWS gives that figure and the price per model unit through the account manager.

What have we seen in Bedrock and EDP reviews since 2024?

Across roughly 20 to 30 AWS Bedrock and EDP reviews in 2024 to 2025, Bedrock token spend was the fastest growing and least governed line in the AWS bill. Most cost advice starts with the rate card: pick the cheaper model, watch output tokens, consider Provisioned Throughput. All of that is sound, but it is not where the money went.

The governance gap sat upstream of every pricing decision. The question was whether the spend was visible to the people negotiating the contract at all, and it routinely was not. Bursty workloads, by contrast, were usually left on demand correctly.

The spend that would have moved the company up a discount band was already happening. It simply was not in the spreadsheet that set the commitment.

Two mechanical traps sit underneath all of it. The first is output pricing, covered above. The second is geography: Bedrock pricing varies by region and some models are available in only one region, so a data residency requirement becomes a rate difference.

An analytics dashboard with charts open on a laptop screen
Bedrock usage shows up in CloudWatch per model long before it shows up in a finance forecast, which is why the commitment conversation should start from engineering data.

Routing adds a related choice. AWS says global cross region inference saves approximately 10 percent on input and output tokens against geographic cross region inference, but requests may be processed in any supported commercial region. Your compliance team has to approve that trade before engineering takes the discount.

What should you ask AWS for in a Bedrock negotiation?

Ask for terms that protect the price of the models you depend on, and make sure every Bedrock dollar counts toward the commitment. Expect the account team to lead with credits and adoption programs, which help but expire.

What the account team will say, and what to say back

  • "Commit to your full AI roadmap and we can improve the discount." Size the Bedrock line from your own CloudWatch and billing data, commit only the base that survives if one project stalls, and ask for a ramp instead of a flat annual figure.
  • "We can add AI credits to this renewal." Ask how much of the credit value you can use before it expires, and what discount the same value would buy instead.
  • "Provisioned Throughput will guarantee your performance." Ask for the model unit throughput for your model in writing, then compare it with your measured tokens per minute before signing a term.
  • "The newest model is where the roadmap is going." Agree to test it, then show the evaluation results for the cheaper candidates alongside it.

Contract wording to ask for

  • Bedrock in qualifying spend. Name Bedrock, including Marketplace billed models, in the qualifying spend definition, so model usage retires the commitment in full.
  • Model version and lifecycle notice. Pin the model versions you run in production and ask for notice longer than the standard Legacy period, because AWS deprecates older versions on a rolling schedule.
  • Extended access pricing. AWS states that customers on Provisioned Throughput or private pricing keep their current terms during extended access. Ask for the same price hold on your on demand models, or migration credits if one enters extended access at a higher rate.
  • Regional availability. Record the regions your residency rules require and the models available there, so a withdrawal does not force an unpriced move.
  • Throughput flexibility. Ask for the right to move a Provisioned Throughput commitment to a newer model version within the same term.

What to do next

  1. Now, in engineering, whatever the renewal date. Constrain output length where quality allows, turn on prompt caching for repeated context, and move tolerant jobs to batch or Flex.
  2. 12 months before the EDP renewal. Add Bedrock to the commitment forecast with its growth curve, and check whether the corrected total crosses a threshold the old forecast did not reach.
  3. 9 months before. Run three to five models over the same prompts for each production workload and compare quality, latency and cost. Put the review on an annual cycle.
  4. 6 months before. Segment workloads by sustained use, reserve throughput only for the portion that clears the breakeven, and leave the bursty remainder on demand.
  5. 3 months before. Confirm how Marketplace billed models count toward the commitment, and price any credit offer against the discount it replaces.
  6. Before signature. Pin model versions and regional availability in the contract so a rolling deprecation does not force an unpriced migration. Our AWS advisory team can rebuild the commitment model with you, and the wider library sits in the AWS knowledge hub.
When to bring in help

Is an AWS commitment coming up? Our AWS EDP negotiation team sizes it to your usage and works only for buyers.

Frequently asked questions

How does AWS Bedrock pricing work?

Bedrock has three commercial shapes. On demand charges per million tokens with separate input and output rates, Provisioned Throughput reserves model units billed hourly on one or six month terms, and customization adds a one off training charge. Batch, Flex and Priority then adjust the on demand rate per request.

Is Bedrock usage included in the AWS EDP commitment?

Yes, and it is the point buyers miss most often. Across our reviews, excluding Bedrock from the forecast left 10 to 25 percent of negotiating room unused. Check separately how models billed through AWS Marketplace are treated in your qualifying spend definition.

How much does model choice matter for Bedrock cost?

It is the largest cost variable you control. Tested models of near identical task quality differed by 5 to 15 times in cost, and on raw list rates a blended workload can differ by close to nineteen times between premium reasoning and lightweight models.

Why do Bedrock output tokens cost more than input tokens?

Output tokens are generated one at a time and cost the provider more compute, so on most commercial models they run three to five times the input rate. Set maximum output lengths per use case, ask for shorter structured answers where users accept them, and cache long prompts that repeat.

Should we move Bedrock workloads to Provisioned Throughput?

Usually not. Breakeven sits around forty percent use of the reserved capacity, and a typical enterprise assistant is busy in office hours and quiet overnight and at weekends. Provisioned Throughput suits sustained production inference at high volume and is required for most custom models.

Then why did steady workloads overpay on demand?

Because the choice has to be made per workload, and for the steady ones it was never made. They ran 30 to 60 percent above what Provisioned Throughput would have cost. Plot tokens per minute over a month and commit only the stable base.

How much does the Provisioned Throughput term length matter?

The six month commitment runs roughly forty percent below the one month rate. That makes term the biggest price difference once a workload qualifies, but a longer term also locks the model version for six months, so check its lifecycle status first.

What EDP discount should we expect?

AWS publishes no EDP rate card, so these are the bands we observe. Discounts step by commitment and do not exceed 20 percent: roughly 5 to 10 percent between $1m and $5m, 10 to 15 percent between $5m and $25m, and 15 to 20 percent above $25m. Term length and ramps change where you land inside a band.

Is the AWS Marketplace contribution to the EDP a discount?

No. The cap on Marketplace spend counting toward the commitment is a separate mechanism from the discount ladder. Use it to reach a threshold, and model it apart from the discount rate so the business case does not overstate the saving.

What should we pin in a Bedrock contract?

The model versions you run in production and the regions they must run in. Most Bedrock models get a 6 month Legacy period before end of life, some only 45 days, and prices can change during extended access. Written notice and price protection cost AWS little to give.

Newsletter
Licensing news that changes what you pay

One email a week on vendor price moves, audit activity and what worked in recent renewals.

Subscribe
Vendor Shield
An advisor on call for every vendor conversation

Always on advisory for renewals, audits and contract questions across your software vendors.

Explore Vendor Shield
Advisory White Paper

Get the AWS EDP negotiation guide.

The commitment sizing worksheet, ramp structures, Marketplace inclusion wording and benchmark bands by commitment size and term.

Gated with a work email on the download page. No sales follow up you did not ask for.

Get the White Paper →
We never share your details with vendors.

AWS licensing news, once a week.

Price changes, audit activity and what worked in recent renewals. No vendor spin.