The 99.9 is the least valuable number in your AI contract
The Azure OpenAI availability SLA answers exactly one question: did the endpoint return a response. The incidents that actually reach an executive are slow responses on a service that never stopped answering, and no percentage on the availability line reaches those at all. There is a separate, published latency SLA almost nobody quotes back across the table, a support decision that turns entirely on one severity tier, and a claim deadline that quietly kills real credits.
Prepared by Redress Compliance · August 9, 2026 · Microsoft advisory. Based on 12 to 18 Azure OpenAI commitment reviews run 2024 to 2025.
Executive summary
The credit ladder is short, it pays in discount not cash, and 43.2 minutes is the trigger.
Below 99.9 percent monthly uptime pays 10 percent of the affected service charge for that month, below 99 percent pays 25, and a 30-day month is 43,200 minutes, so nothing is owed until downtime passes 43.2 minutes and the 25 percent rung starts at 7 hours 12 minutes.
A 95-minute incident on a deployment billing 180,000 dollars a month is 99.78 percent up, the first rung, 18,000 dollars of credit against a future invoice; run the same outage eight hours and it moves to 45,000.
That is the top of Microsoft's financial exposure to your worst AI day, so weigh it against what a dead assistant costs the business before you fight for another nine.
Azure claims run two months, and every other Microsoft service is a full month earlier, so one runbook gets one of them wrong.
The Claims section gives Azure until two months after the end of the billing month; every other Microsoft online service must be claimed by the end of the following calendar month, so a February incident on M365 is dead on April 1 while the same incident on Azure OpenAI still has April.
The claim must carry four things, a detailed description, the time and duration, the number and location of affected users, and a description of your attempts to resolve it at the time, and the fourth kills retrospective claims: if nobody opened a support case while the service was down.
The claim fails on paperwork rather than facts.
Open the case during the outage even when you already know it is Microsoft's problem.
There is a latency SLA, it is a floor, and it is the only performance commitment Microsoft has already put in writing.
Microsoft committed to 99 percent on token generation and publishes per-model numbers in a Latency Target column, on gpt-4o a floor of 25 tokens per second, provisioned deployments only.
Three things shrink it: it is a floor on output token rate, not a response time; it is measured as p50 on a per five-minute basis, so a bad five minutes inside a good hour vanishes.
And it lives on provisioned deployments, because on Standard the metrics that would evidence it are not emitted at all.
A 600-token answer at 25 tokens per second is 24 seconds to the last token, well inside Microsoft's own floor, so if a product owner promised a two-second answer, the gap is a number the architects should see before the reservation is signed.
Pin the target to the Effective Date and attach a credit to it.
The support decision is one severity tier, and Unified is priced on your spend, so the AI line drags the bill up behind it. Azure support lists at 29, 100 and 1,000 dollars a month for Developer, Standard and Professional Direct, and the SLA is not support: severity response comes from the plan.
A model that stopped answering is a Sev A that Standard already answers in an hour.
The incident you will actually file is degraded latency on a service still returning 200s, a Sev B, and the 900 dollars a month between Standard and Professional Direct buys exactly one thing, half the wait, four hours down to two, on the only severity your AI incidents will ever qualify for.
Unified runs roughly 8 to 10 percent of Azure consumption, so a 2.16 million dollar a year deployment adds 172,800 to 216,000 to a Unified bill before anyone opens a case, against Professional Direct at 12,000 flat.
What the SLA covers versus what production needs
| Risk | In the SLA? | Where protection actually comes from |
|---|---|---|
| Service unavailable | Yes, 99.9 percent | Credit claim after breach, 10 to 25 percent |
| Slow responses, latency spikes | Partly, provisioned only | The published Latency Target, plus a drafted response-time clause |
| Throttling at high load | No | PTU reservation sizing on measured tokens |
| Model quality regression | No | Deployment pinning and evaluation gates |
| Model retirement | No | Migration planning and contract notice terms |
Time to first token, sustained throughput under burst, and output quality sit outside both commitments, and those are the failure modes that reach an executive.
The availability number counts responses, not seconds, so the availability percentage will not move in negotiation, but a response-time service level can, because Microsoft has already published the number and already calls it an SLA.
What is missing is a remedy attached to it: ask the account team to point at the credit ladder for the latency SLA, and if nobody can, that gap is the clause.
Draft a Model Response Time Service Level that holds a Token Rate Percentage of not less than 99 percent, measured by the AzureOpenAITokenPerSecond metric against the Latency Target published as of the Effective Date, with a 10 percent credit on failure.
Pin the table to the Effective Date, because an unpinned reference lets a model retirement reset your service level downward with your signature already on it, and the latency-sensitive path has to sit on provisioned capacity for the metrics to exist at all.
The pricing detail sits in the Azure OpenAI pricing guide.
Buying predictable performance with PTUs
- Provisioned throughput is the latency instrument: PTUs reserve model processing capacity with consistent latency, where Standard pay-as-you-go shares capacity and absorbs the noisy-neighbor problem. Blend the modes, PTU for the latency-sensitive core, pay-as-you-go for batch and overflow.
- Size on measured tokens, not launch forecasts: base PTU counts on observed tokens per minute from a pilot against the published input TPM per PTU, measured with AzureOpenAIProvisionedManagedUtilizationV2. A measured 150,000 input tokens per minute on gpt-4o is 60 PTU on Global, buyable exactly on a 5 PTU increment.
- Data residency can buy 40 PTU of nothing: the same 60 PTU of need on Regional Provisioned has a minimum of 50 and an increment of 50, so the smallest size that covers you is 100 PTU, 67 percent more capacity than the workload consumes.
- Model choice moves it 3.6x: that 150,000 tokens per minute needs 31.6 PTU on gpt-5 rounding to 35, but 125 PTU on gpt-5.5, same workload, same region, roughly 3.6 times the capacity purchase for what reads on a release note like a version bump, so rerun the arithmetic before every model migration.
- A reservation does not reserve capacity and cannot move between types: you deploy first and buy the reservation second, deployed PTUs above the reserved quantity bill hourly, and reservations cannot be exchanged between Global, Data Zone and Regional, so changing type means cancelling, which can carry an early termination fee. Target above 60 percent sustained utilization before adding capacity.
The Azure OpenAI commitment playbook
The SLA remedies, the latency clause, the PTU sizing math, and the support-tier decision, worked end to end.
Get the white paper →Which support plan an AI workload actually needs
The SLA is not support: severity response times, escalation paths and engineering access come from your Azure support plan, purchased separately or wrapped into a Unified agreement.
Basic is included with no technical support, Developer at 29 dollars a month covers only Sev C in 8 business hours, Standard at 100 answers a Sev A in one hour and a Sev B in four, and Professional Direct at 1,000 answers a Sev B in two.
Read that against what actually happens to an AI workload: a model that has stopped answering is a Sev A, and Standard already gives you a one-hour response on it, so the incident you will really file is degraded latency on a service still returning 200s, which is a Sev B.
And the 900 dollars a month between Standard and Professional Direct buys exactly half the wait on the only severity your AI incidents will ever qualify for.
That is the whole decision, and it is usually made by someone who never read the severity table.
Unified is the trap underneath, because Microsoft does not publish its rates and prices the model as a percentage of your annual Microsoft spend by category, not a fee for the support you consume.
Roughly 8 to 10 percent of online services spend including Azure consumption: a deployment billing 180,000 dollars a month is 2.16 million a year.
Adding roughly 172,800 to 216,000 to a Unified bill before anyone opens a case, while Professional Direct is 12,000 a year flat and does not care how big the AI estate gets.
When the AI line is the thing growing fastest, the arithmetic for pricing Azure separately against Professional Direct is worth putting on paper before the support renewal, not after. Copilot or M365 support does not cover Azure OpenAI, which is an Azure service under Azure support terms.
The plan detail sits in the Azure OpenAI negotiation guide.
- Percentile standing for your exact deal size and industry, from real closed transactions
- Scenario simulation before the call: test alternative terms and see the financial impact of each
- A negotiation playbook, talking points, and a two page executive brief on day one
What we saw across Azure OpenAI commitment reviews, 2024 to 2025
Fredrik Filipsson ran between 12 and 18 Azure OpenAI commitment reviews across 2024 and 2025, and every one spent its argument on the 99.9 number while none of the incidents that reached a steering committee afterwards had anything to do with it:
Where provisioned throughput sized against launch forecasts was running by month three, so a third to a half of reserved AI capacity was paid headroom.
Of committed AI spend moved by negotiating around the SLA: PTU reservation pricing, term flexibility, and migration support, not the availability percentage.
Not one incident in that file that caused measurable business damage was an availability breach; they were latency collapses on a service that kept answering, and no percentage on the availability line reaches those.
In two cases the Azure claim deadline expired before anyone had assembled the four items the Claims section requires, so the credit was real and went unclaimed.
An extra nine costs Microsoft very little to concede and pays a credit on an outage you will probably never have, so put the capital where it lands: a credit remedy attached to the published Latency Target, that target pinned to the Effective Date, the Sev B response number rather than the Sev A one.
And model retirement notice.
Build the evidence on four Azure Monitor metric IDs under Microsoft.CognitiveServices/accounts, AzureOpenAITimeToResponse, AzureOpenAINormalizedTTFTInMS for time to first byte, AzureOpenAINormalizedTBTInMS for time between tokens, and AzureOpenAITTLTInMS for time to last byte.
Rather than the legacy Latency metric most dashboards still chart, and note that two of them plus Tokens Per Second are not emitted for Standard deployments at all.
Tie AI spend into MACC drawdown so it earns your existing discount structure, and cap the reservation term while model economics are falling. The comparison against buying direct sits in the Azure OpenAI versus direct comparison, and the EA framing in the EA renewal playbook.
Your first five moves
- Stop pushing the availability percentage and attach a credit remedy to the published Latency Target instead, pinned to the Effective Date so a model retirement cannot reset it downward.
- Size PTUs on measured tokens from a pilot, targeting above 60 percent sustained utilization, and rerun the arithmetic before every model migration because the migration is the price change.
- Buy the Sev B number, not the Sev A: Professional Direct's two-hour Sev B against Standard's four is the only support difference an AI incident will ever feel.
- Price Azure separately against Professional Direct before a Unified renewal, because Unified charges a percentage of the fast-growing AI line for support you may not consume.
- Rehearse the SLA claim path and assign an owner before go-live, opening a support case during any outage, because the fourth claim item kills retrospective claims. The Microsoft practice runs it buyer side.
Frequently asked questions
What does the Azure OpenAI SLA actually cover?
Availability only: a 99.9 percent monthly uptime commitment that answers exactly one question, did the endpoint return a response. Time to first token, sustained throughput under burst, and output quality all sit outside it.
Microsoft also carries a separate 99 percent latency SLA for token generation on provisioned deployments, but that is a floor on output token rate measured p50 per five minutes, not a response-time guarantee.
The failure modes that reach an executive, latency collapses on a service that kept answering, are not covered by the availability number.
How much does an Azure OpenAI outage pay?
Ten percent of the affected service charge for the month below 99.9 percent uptime, or 25 percent below 99 percent, always as a credit against a future invoice, never cash, and never against the whole Azure bill.
In a 30-day month nothing is owed until downtime passes 43.2 minutes, and the 25 percent rung starts at 7 hours 12 minutes. A 95-minute incident on a deployment billing 180,000 dollars a month is worth 18,000; eight hours is worth 45,000.
That is the ceiling on Microsoft's exposure to your worst AI day.
What is the deadline to claim an Azure OpenAI service credit?
Two months after the end of the billing month in which the incident occurred, which is a full month longer than every other Microsoft online service, so a single incident runbook across the estate gets one of them wrong.
The claim must carry four things: a detailed description, the time and duration, the number and location of affected users, and a description of your attempts to resolve it at the time.
The last kills retrospective claims, so open a support case during the outage even when you know it is Microsoft's problem.
Is there a latency SLA for Azure OpenAI?
Yes, a 99 percent latency SLA for token generation announced in November 2024, with per-model numbers in a Latency Target column, for example a floor of 25 tokens per second on gpt-4o.
Three things shrink it: it is a floor on output token rate, not a response time, so a 600-token answer can take 24 seconds and still comply; it is measured p50 per five minutes; and it exists only on provisioned deployments, because on Standard the evidencing metrics are not emitted.
It is the only performance commitment Microsoft has already put in writing, so quote it back and attach a credit.
Which Azure support plan does an AI workload need?
Standard at 100 dollars a month covers the Sev A of a model that stopped answering in one hour, but the incident you will actually file is degraded latency on a service still returning 200s, a Sev B.
Standard answers a Sev B in four hours and Professional Direct at 1,000 a month in two, so the 900 dollars a month between them buys exactly half the wait on the only severity your AI incidents qualify for.
The SLA is not support, and Copilot or M365 support does not cover Azure OpenAI, which runs under Azure support terms.
Should Azure OpenAI support go through Microsoft Unified?
Weigh it carefully, because Unified is priced as a percentage of your annual Microsoft spend, roughly 8 to 10 percent of online services spend including Azure consumption, not a fee for the support you consume.
A deployment billing 180,000 dollars a month is 2.16 million a year, adding roughly 172,800 to 216,000 to a Unified bill before anyone opens a case, while Professional Direct is 12,000 a year flat regardless of estate size.
When the AI line is the fastest-growing spend, price Azure separately against Professional Direct before the support renewal.