The PTU reservation is a capacity contract that happens to carry a discount
Azure OpenAI is the fastest growing line in many Microsoft accounts, running at 8 to 22 percent of the total Azure invoice on AI heavy estates and growing 60 to 140 percent year on year. It sits inside the MACC envelope, which means the AI line is negotiated against the wider Azure commitment rather than beside it. The clauses set the price ceiling and the operational guarantees for the next three years, and the model catalog turns over faster than the contract does.
Prepared by Redress Compliance · August 10, 2026 · Microsoft advisory. Based on the Azure AI commitment review file, 2024 to 2025.
Executive summary
Reserved PTU prices 18 to 34 percent below pay as you go, and the reservation is also the operational SLA.
Provisioned Throughput Units reserve dedicated throughput at a flat hourly rate, sized to a target tokens per minute with a minimum deployment per model, typically 50 to 100 PTU on flagship families. The discount is only half the value.
On heavily used regions and flagship models the pay as you go tier returns throttling under peak demand, while reserved PTU never queues against shared capacity, so the reservation is what converts a best effort service into a predictable one.
The crossover sits at the load curve, not at a single utilization threshold.
Pay as you go bills per thousand input and output tokens at a rate that varies by model family, context length, and region, and it is the right vehicle for spiky, exploratory, and low volume work where capacity guarantees do not justify a reservation.
Reserved PTU fits steady, high throughput, latency sensitive workloads. The route that wins in practice is a split: anchor the steady throughput on reserved PTU and handle the spikes on pay as you go, rather than forcing one vehicle across the whole estate.
The model catalog turns over every six to nine months, so lifecycle protection has to be in writing. Models are deprecated, replaced, or repriced with limited notice, and standard notice runs 6 to 12 months.
Three protections keep an architecture stable across a three year term: a minimum 12 month notice on the deprecation of any production model, a substitution right to a replacement at no worse price or performance.
And a price hold billing the replacement at the deprecated model rate through a transition window.
Without them, the vendor's release cadence sets your migration schedule.
Bundle the AI line into MACC, because it can move the wider Azure discount band by a tier.
Azure OpenAI consumption counts toward the Microsoft Azure Consumption Commitment on both reserved PTU and pay as you go, though marketplace billed third party model offerings may not, so verify each line at quote time.
Used deliberately, the AI commitment clears the threshold for the next discount tier on the whole Azure estate, which is worth more than negotiating the token rate in isolation.
Fine tuning stays a separate budget line: the training run, the hourly deployment fee per hosted model, and a higher inference rate each bill separately.
PTU against pay as you go
| Dimension | Pay as you go | Provisioned Throughput Units |
|---|---|---|
| Pricing unit | Per 1,000 input and output tokens | Per PTU per hour, flat |
| Commitment | None | Hourly, monthly, one year, or three year reserved |
| Capacity guarantee | Best effort | Reserved throughput, no queueing |
| Discount versus list | 0 percent | 18 to 34 percent on reserved one and three year |
| Minimum unit | One request | Per model minimum, typically 50 to 100 PTU |
| Best fit workload | Spiky, exploratory, low volume | Steady high throughput, latency sensitive |
| Region availability | Wide | Narrower, model and region dependent |
The reservation is more than a discount, and buyers who price it as one under value it.
On heavily used regions and flagship models the shared pay as you go tier returns throttling under peak demand, which means an application that must respond has no contractual claim on capacity at the moment it matters most.
Reserved PTU does not queue against shared capacity, so the reservation functions as the operational SLA for the workload. Price the reliability alongside the rate, because for a mission critical application the throttling risk is usually the larger of the two numbers.
The wider Azure commitment mechanics sit in the Azure MACC negotiation guide.
Five clauses every buyer needs
| Clause | What it does | What it protects |
|---|---|---|
| Reserved PTU capacity guarantee | Locks dedicated throughput | Operational SLA and discount |
| Model lifecycle protection | Notice and substitution rights on deprecation | Architecture continuity |
| Data residency lock | Specifies the region for inference and logging | Regulatory compliance |
| Data use opt out | Confirms data is not used for training | IP protection and privacy |
| Annual price cap | Bounds annual PTU price increases | Three year price predictability |
The Microsoft EA renewal playbook
The Azure and Azure OpenAI renewal cycle: PTU benchmarks, MACC commit arithmetic, model lifecycle clauses, and the residual clause checklist.
Get the white paper →The commercial levers and the three routes that win
Six levers move the AI line, and only one of them is the token rate. The reserved PTU term matters first, because one year and three year reservations sit at different discount bands and the term is the cheapest thing to lengthen once the workload is genuinely steady.
Model mix comes next: negotiate price separately on the GPT family, on embeddings, on fine tuned models, and on image models, rather than accepting a single blended figure that hides which lines you actually consume.
Region commitment carries a capacity premium on multi region deployments, so decide deliberately rather than defaulting to breadth. Volume tier moves the discretionary discount band as monthly token volume rises.
The MACC anchor is the lever most buyers under use, because pulling AI commitment into the pool can clear the threshold for the next discount tier across the entire Azure estate.
And promotional credit is routinely available on multi year Azure commits, which is worth asking for explicitly rather than waiting to be offered. Three buyer routes recur. Reserved PTU heavy, where steady throughput anchors on reservations and spikes run on pay as you go.
MACC bundled, where the AI line is used to accelerate the wider Azure discount tier. And model agnostic, where the architecture abstracts the model and the contract keeps the option to switch open, which is the route that ages best given a catalog that turns over twice a term.
The comparable vehicle on the other hyperscaler sits in the Bedrock pricing guide, and the seat side of Microsoft AI in Copilot pricing.
- Percentile standing for your exact deal size and industry, from real closed transactions
- Scenario simulation before the call: test alternative terms and see the financial impact of each
- A negotiation playbook, talking points, and a two page executive brief on day one
Why the AI contract carries weight
The Azure OpenAI line is the fastest growing item in many Microsoft accounts, and three characteristics make its contract terms disproportionately valuable relative to its current share of spend:
Year on year growth in AI workloads on heavy use estates, which means today's minor line is next term's major one.
The reserved PTU discount against pay as you go, before counting the value of guaranteed capacity on flagship models in busy regions.
Scale is the first: AI workloads grow at 60 to 140 percent year on year on heavy use estates, so terms agreed on a small line govern a large one within the term.
Concentration is the second: Azure OpenAI is the single largest AI line for most Microsoft customers, which removes the diversification that usually limits the damage of a weak clause.
Capacity scarcity is the third: model availability and regional capacity are not guaranteed without contract language, so the risk that bites is operational rather than financial.
The buyer side move is to pull the trailing six month token baseline by model and region, compute the PTU equivalent, score workload predictability to set the reserved and pay as you go split, then build a three year envelope and draft the five clauses before the next MACC conversation opens.
Your first five moves
- Pull the token consumption baseline by model and region for the trailing six months, then compute the PTU equivalent against the per model minimums of 50 to 100 PTU.
- Score workload predictability and set the split, anchoring steady throughput on reserved PTU at 18 to 34 percent below pay as you go and leaving the spikes uncommitted.
- Map the model dependency tree application by model by region, and inventory the data residency obligations, because both become clauses rather than settings.
- Draft the five clauses before the commercial conversation: capacity guarantee, model lifecycle with 12 month notice and a substitution right, residency lock, data use opt out, and an annual price cap.
- Bundle the AI line into MACC to clear the next Azure tier, verifying which lines qualify, since marketplace billed third party models may not. Open the file 6 months before the MACC renewal. The Microsoft practice runs the envelope and the negotiation with you.
Frequently asked questions
Is Azure OpenAI cheaper than calling OpenAI directly?
Per token they price similarly on most flagship models. The commercial advantage on Azure sits elsewhere: the MACC bundle, the reserved PTU discount, the data residency commitment, and the enterprise compliance posture.
The buyer side decision is rarely about the token unit price alone, which is why negotiating the token rate in isolation leaves the larger levers untouched.
What is a Provisioned Throughput Unit?
PTU reserves dedicated throughput at a flat hourly rate, sized to a target tokens per minute with a minimum deployment per model. Reserved PTU at one or three year terms carries the deepest discount, 18 to 34 percent below pay as you go, and the strongest capacity guarantee.
Flagship GPT family models typically require 50 to 100 PTU as the minimum deployment per model in a region.
Why is reserved PTU also a capacity contract?
Because on heavily used regions and flagship models the shared pay as you go tier returns throttling under peak demand, and an application with no reservation has no contractual claim on capacity at the moment it matters.
Reserved PTU never queues against shared capacity, so the reservation functions as the operational SLA rather than only as a discount.
How does Azure OpenAI work with the MACC commitment?
Azure OpenAI consumption counts toward the Microsoft Azure Consumption Commitment, and both reserved PTU and pay as you go qualify. Marketplace billed third party model offerings may not, so verify each line at quote time.
Bundling the AI commitment into the MACC can move the wider Azure discount band by a full tier, which is usually worth more than the AI line's own rate.
What is the model deprecation notice period?
Standard notice runs 6 to 12 months depending on the model, and the catalog turns over every six to nine months.
Enterprise buyers should negotiate a minimum 12 month notice on any production model, a substitution right to a replacement at no worse price or performance, and a price hold billing the replacement at the deprecated model rate through a transition window.
Is Azure OpenAI data used for training?
No. The default service does not use customer prompts or completions to train Microsoft or OpenAI models, and the position is documented in the product terms.
Buyers in regulated industries should still confirm the data use opt out in writing in the order form and specify the data centre region for inference and any logging, because residency is a clause rather than a setting.
How should fine tuning be budgeted?
As a separate line. Fine tuning bills for the training run and for deployment time, deployment carries an hourly fee per hosted fine tuned model, and inference on a fine tuned model bills at a higher token rate than the base model.
Buyers planning extensive fine tuning should price the deployment hours and the inference rate independently rather than folding them into the base envelope.