Rack mounted server hardware with rows of status lights
Azure OpenAI SLA

The Azure OpenAI SLA and support plans. What the 99.9 percent pays, and what it never covers.

The availability credit ladder, the Azure claim deadline, the published latency SLA on provisioned deployments, PTU sizing and the support plan choice for AI workloads.

Contact Us Microsoft Advisory
500+Enterprise clients
$2B+Under advisory
PublishedMay 21, 2026UpdatedSeptember 23, 2026
ContentsKey takeawaysWhat the SLA coversWhat an outage paysClaim deadline and contentsThe latency SLASizing PTUsChoosing a support planWhat we have seenWhat to ask Microsoft forPilot versus productionWhat to do nextFAQ

The Azure OpenAI SLA only measures whether the endpoint answered. The incidents that hurt are slow answers from a service that stayed up, so the latency target, the claim rules and the Sev B response time matter more than the 99.9.

Key takeaways
  • The credit ladder is short. Below 99.9 percent uptime Microsoft credits 10 percent of the month's Azure OpenAI charge, and below 99 percent it credits 25 percent, never as cash.
  • Nothing is owed for the first 43.2 minutes. In a 30 day month the 10 percent rung starts after 43.2 minutes of downtime and the 25 percent rung after 7 hours 12 minutes.
  • Azure claims run two months. Other Microsoft online services must be claimed a month earlier, and a claim without a support case opened during the outage is likely to fail.
  • A latency SLA exists on provisioned deployments. Microsoft publishes a Latency Target per model, 25 tokens per second on gpt-4o, measured as a p50 per five minutes.
  • Model choice changes the PTU bill. The same 150,000 input tokens per minute needs 35 PTU on gpt-5 and 125 PTU on gpt-5.5.
  • Support turns on Sev B. Professional Direct costs $900 a month more than Standard and halves the Sev B wait, the severity a latency incident gets filed under.

The Azure OpenAI availability SLA promises 99.9 percent monthly uptime and pays a service credit of 10 to 25 percent of that month's charge when Microsoft misses it. It measures one thing: whether the endpoint returned a response. It says nothing about how long the response took.

That gap matters because the AI incidents that reach a steering committee are rarely outages. They are slow answers from a service that kept returning responses the whole time.

What does the Azure OpenAI SLA actually cover?

The availability SLA covers one risk: the service not answering at all. Time to first token, sustained throughput under burst load and the quality of what the model returns all sit outside it. Microsoft now documents the service as Azure OpenAI in Microsoft Foundry Models, but the SLA terms work the same way.

Map each production risk to the place where protection has to come from. Only the first row is covered by the 99.9 percent number.

What the SLA covers against what a production AI workload needs
RiskIn the SLA?Where protection comes from
Service unavailableYes, 99.9 percentA credit claim after the breach, 10 or 25 percent of the monthly charge
Slow responses and latency spikesPartly, provisioned deployments onlyThe published Latency Target, plus a response time clause you draft
Throttling at high loadNoPTU reservation sizing on measured tokens
Model quality regressionNoPinning the deployment to a model version, and evaluation gates before any upgrade
Model retirementNoMigration planning and notice terms in the contract

The rest of this page follows those rows. The Azure OpenAI SLA coverage review goes further into the service level text itself.

Watch the briefingResearch briefing · 4:33

How to Negotiate with OpenAI and Anthropic: The Vendors With Nobody to Call

How much does Microsoft pay when Azure OpenAI misses the SLA?

Below 99.9 percent monthly uptime, Microsoft owes a credit of 10 percent of the affected service charge for that month. Below 99 percent, the credit rises to 25 percent. There is no third rung.

  • Credit, never cash. The money comes back as a discount on a future invoice.
  • Affected service only. The percentage applies to the Azure OpenAI charge for the month, never to your whole Azure bill.
  • A threshold before anything is owed. A 30 day month has 43,200 minutes, so 0.1 percent of it is 43.2 minutes of downtime. Nothing is payable below that.
  • The upper rung starts late. The 25 percent credit needs 1 percent downtime, which is 7 hours 12 minutes in a 30 day month.

Worked example: $180,000 of monthly Azure OpenAI charges

Say your Azure OpenAI deployment bills $180,000 a month. The table shows what four outages of different length would return in a 30 day month.

Service credit on a $180,000 monthly Azure OpenAI charge (30 day month)
Outage lengthMonthly uptimeCredit rungCredit on the next invoice
40 minutes99.91 percentNone$0
95 minutes99.78 percent10 percent$18,000
6 hours99.17 percent10 percent$18,000
8 hours98.89 percent25 percent$45,000

A 95 minute incident and a six hour incident pay the same $18,000. The $45,000 for an eight hour outage is the most Microsoft stands to lose on your worst AI day. Weigh that against what a dead customer assistant costs the business per hour.

Why we would not spend negotiating time on another nine

The usual advice is to push Microsoft from 99.9 to 99.95 percent or higher. We disagree. An extra nine costs Microsoft very little to concede, because it pays a credit on an outage you will probably never have.

A higher figure only changes when the same small credit is owed, and the table above shows how little even an eight hour outage returns. Spend the negotiating time on a credit tied to the published Latency Target, the Sev B response time and model retirement notice, which are the terms our review findings below point to.

Free white paper

Azure OpenAI commitment guide

SLA remedies, latency clause wording, PTU sizing and the support plan choice in one download.

Get the white paper →

What is the deadline to claim an Azure OpenAI service credit?

For Azure services, including Azure OpenAI, Microsoft must receive the claim within two months after the end of the billing month in which the incident happened. Every other Microsoft online service has to be claimed by the end of the following calendar month.

That one month difference trips up teams that run a single incident runbook for all of Microsoft. Take an incident in February. On Microsoft 365 the claim window closes on April 1. The same incident on Azure OpenAI can still be claimed during April. A runbook written for one of those services gets the other wrong.

What the claim has to contain

The Claims section of the SLA asks for four items. Microsoft can reject a claim that misses any of them on paperwork, whatever the facts were.

  1. A detailed description of the incident.
  2. The time and duration of the downtime.
  3. The number and location of affected users.
  4. A description of your attempts to resolve the incident at the time it occurred.

The fourth item defeats claims assembled after the event. If no support case was opened while the service was down, there is no record of an attempt to resolve it. Open a case during the outage even when the Azure status page already shows the problem is on Microsoft's side.

Who should own the claim

Name one person, usually in cloud operations or vendor management, who owns SLA claims for Azure. Their checklist for each incident:

  • The support case number opened during the outage.
  • The incident tracking ID from Azure Service Health.
  • Azure Monitor exports for the incident window.
  • The count and location of affected users, taken from your own logs.
  • The date the claim window closes, entered in the team calendar on the day of the incident.

The claim itself is a support request in the Azure portal, filed with the issue type Billing and the problem type Refund Request. Microsoft aims to process claims within 45 days, and an approved credit can never exceed that month's fee for the service.

Rehearse the claim once before go live. In two reviews we saw the Azure deadline expire before anyone had gathered the four items, so a real credit went unclaimed.

Is there a latency SLA for Azure OpenAI?

Yes. In November 2024 Microsoft announced a 99 percent latency SLA for token generation on provisioned deployments. It publishes a Latency Target for each model, and it is the only performance commitment Microsoft has already put in writing. Almost no buyer quotes it back during the negotiation.

Published Latency Targets for provisioned deployments (selected models)
ModelLatency TargetTime to generate a 600 token answer at the floor
gpt-4o99 percent above 25 tokens per second24 seconds
gpt-4o-mini99 percent above 33 tokens per secondAbout 18 seconds
gpt-4.199 percent above 80 tokens per second7.5 seconds
gpt-5, gpt-5.1 and gpt-5.599 percent above 50 tokens per second12 seconds

Three limits that shrink the latency SLA

  • It is a floor on output token rate. It does not promise a response time. A 600 token answer at 25 tokens per second takes 24 seconds to the last token and still complies.
  • It is measured as p50 on a per five minute basis. A bad five minutes inside a good hour disappears into the median.
  • It applies to provisioned deployments only. On Standard deployments, the metrics that would evidence a breach are not emitted at all.

If a product owner promised users a two second answer, show the architects the 24 second figure before the reservation is signed. The gap between the promise and Microsoft's own floor is something your design has to cover, through shorter answers, streaming or a faster model.

A developer at a desk watching several monitoring dashboards
Latency evidence has to be collected while an incident is happening. Azure Monitor keeps platform metrics for 93 days by default, so export the incident window to a Log Analytics workspace or storage account when you open the case.

How to evidence latency with Azure Monitor

Build your latency evidence on the Azure OpenAI metrics under Microsoft.CognitiveServices/accounts. Most dashboards still chart the legacy Latency metric, which Microsoft says is not designed for Azure OpenAI and gives misleading results.

  • AzureOpenAITimeToResponse. Time to response for each request.
  • AzureOpenAINormalizedTTFTInMS. Normalized time to first byte.
  • AzureOpenAINormalizedTBTInMS. Time between tokens.
  • AzureOpenAITTLTInMS. Time to last byte.
  • AzureOpenAITokenPerSecond. The token rate that the Latency Target is written against.

Two of the first four, time to response and time between tokens, are not emitted for Standard deployments, and neither is Tokens Per Second. That is why the latency sensitive path has to run on provisioned capacity if you want a latency commitment you can enforce.

How do provisioned throughput units buy predictable performance?

Provisioned throughput units (PTUs) reserve model processing capacity with consistent latency. Standard deployments share capacity with other tenants and absorb the noisy neighbor problem. Most production designs blend the two: PTUs for the latency sensitive core, and Standard pay as you go for batch jobs and overflow.

Size on measured tokens, never on launch forecasts

Base the PTU count on observed tokens per minute from a pilot, divided by the published input tokens per minute (TPM) per PTU for your model and deployment type. Track it afterwards with the AzureOpenAIProvisionedManagedUtilizationV2 metric, and add capacity only once sustained usage runs above 60 percent.

Take a measured 150,000 input tokens per minute. On gpt-4o, Microsoft publishes 2,500 input TPM per PTU, so the workload needs 60 PTU on a Global deployment, where the minimum is 15 and the increment is 5. That is exactly buyable.

The same 150,000 input tokens per minute, sized four ways
Model and deployment typeInput TPM per PTUExact needSmallest buyable size
gpt-4o, Global2,50060 PTU60 PTU
gpt-4o, Regional (minimum 50, increment 50)2,50060 PTU100 PTU
gpt-5, Global4,75031.6 PTU35 PTU
gpt-5.5, Global1,200125 PTU125 PTU

Output tokens add to the count, and the weighting differs by model: Microsoft counts each output token as 4 input tokens on gpt-4o, 8 on gpt-5 and 6 on gpt-5.5. The table counts input only. Run your real input and output mix through Microsoft's sizing calculator before you commit.

What data residency and model upgrades do to the bill

Data residency can leave you paying for 40 PTU the workload never touches. The same 60 PTU of need on Regional Provisioned rounds up to 100 PTU, which is 67 percent more capacity than the workload consumes. Check whether a Data Zone deployment meets your residency requirement before you accept the Regional minimum.

Model choice can change the purchase by a factor of 3.6. The same workload needs 35 PTU on gpt-5 and 125 PTU on gpt-5.5, in the same region, for what a release note presents as a version bump. Rerun the sizing before every model migration, because the migration is a price change.

How the reservation works

  • Deploy first, reserve second. A reservation does not reserve capacity. Create the deployment, confirm the capacity exists, then buy the reservation to cover it.
  • Overage bills hourly. Deployed PTUs above the reserved quantity are charged at the hourly rate.
  • No exchange between types. Reservations cannot be exchanged between Global, Data Zone and Regional. Changing type means cancelling and buying again, and a cancellation can carry an early termination fee.
  • Short terms are available. Microsoft offers one month and one year reservation terms, which helps while model prices are falling.

The PTU and token pricing detail sits in our Azure OpenAI pricing guide.

Which Azure support plan does an Azure OpenAI workload need?

For most AI workloads the choice is between Standard and Professional Direct, and it turns on the Sev B response time. The SLA is not support. Severity response times, escalation paths and engineering access come from your Azure support plan, bought separately or wrapped into a Unified agreement.

Azure support plans and initial response times
PlanList priceSev ASev BSev C
BasicIncludedNo technical support
Developer$29 a monthNot availableNot availableWithin 8 business hours
Standard$100 a monthUnder 1 hourUnder 4 hoursUnder 8 business hours
Professional Direct$1,000 a monthUnder 1 hourUnder 2 hoursUnder 4 business hours

A model that has stopped answering is a Sev A, and Standard already responds to it within an hour. The incident you will actually file is degraded latency on a service still returning HTTP 200s, a Sev B. The $900 a month between Standard and Professional Direct halves that wait, four hours down to two.

That is the decision in full, and it is often made by someone who never read the severity table. One more point for the runbook: Copilot or Microsoft 365 support does not cover Azure OpenAI. It is an Azure service under Azure support terms. Our Azure OpenAI negotiation guide covers the plan terms in the wider deal.

Does Microsoft Unified make sense for Azure OpenAI?

Often it costs more. Microsoft does not publish Unified rates and prices it as a percentage of your annual Microsoft spend by category, which is a different thing from a fee for support you consume. In our experience it runs roughly 8 to 10 percent of online services spend, and Azure consumption counts.

Take the $180,000 a month deployment from the credit example, which is $2.16 million a year. At those rates it adds roughly $172,800 to $216,000 to the Unified bill before anyone opens a case. Professional Direct costs $12,000 a year flat and does not grow with your AI spend.

When the AI line is the fastest growing part of your Microsoft spend, price Azure support separately against Professional Direct before the Unified renewal, not after it.

For the Unified side of that comparison, see choosing a Unified support level and the alternatives to Unified support.

What have we seen in Azure OpenAI commitment reviews in 2024 and 2025?

My co founder Fredrik Filipsson ran between 12 and 18 Azure OpenAI commitment reviews across 2024 and 2025. Every one spent most of its negotiating time on the 99.9 percent number, and none of the incidents that reached a steering committee afterwards had anything to do with it.

  • Oversized reservations. Provisioned throughput sized against launch forecasts was running at 35 to 55 percent of capacity by month three. A third to a half of the reserved AI capacity was paid headroom.
  • Where the savings came from. Buyers moved 15 to 25 percent of committed AI spend by negotiating around the SLA: PTU reservation pricing, term flexibility and migration support. The availability percentage contributed nothing.
  • What hurt the business. Each incident that caused measurable business damage was a latency collapse on a service that kept answering, which no availability percentage reaches.
  • Lost credits. In two cases the Azure claim deadline expired before the four required items were assembled.

Two commercial points belong in the same negotiation. Tie Azure OpenAI spend into your MACC drawdown so it earns your existing discount structure, and cap the reservation term while model economics are falling.

For the MACC side, see our Azure MACC negotiation guide. The Azure OpenAI versus direct OpenAI comparison covers buying outside Microsoft, and the Microsoft EA renewal playbook covers the renewal context.

What should you ask Microsoft for beyond the standard SLA?

The availability percentage will not move in negotiation. A response time service level can, because Microsoft has already published the number and already calls it an SLA. Ask the account team to show you, in writing, the service credit that applies when the Latency Target is missed. If they cannot point to one, that gap is the clause you draft.

Contract wording to ask for

  • A Model Response Time Service Level. A Token Rate Percentage of not less than 99 percent, measured by the AzureOpenAITokenPerSecond metric against the Latency Target, with a 10 percent service credit on failure. This turns the published floor into a remedy.
  • The Latency Target pinned to the Effective Date. Reference the table as published on the Effective Date. With an unpinned reference, a model retirement can reset your service level downward with your signature already on it.
  • Model retirement notice. A stated minimum notice before a pinned model version is retired, plus migration support, so the forced upgrade is not also a surprise capacity purchase.
  • A price hold on added PTUs. Your negotiated reservation rate should apply to any PTUs you add during the term, so a model migration that needs 3.6 times the capacity is at least bought at your price.
  • MACC eligibility in writing. Confirmation that Azure OpenAI consumption and PTU reservations count toward the commitment.

What the account team will say, and what to say back

  • "The 99.9 percent SLA is standard for every customer." Agree, and move on. Tell them the uptime figure is not what you are asking about, and put the latency credit on the table instead.
  • "The latency SLA is already in the service terms." Then ask them to name the credit, and to attach the Latency Target table as of the Effective Date to the order.
  • "Buy the one year reservation now so the capacity is secured." Microsoft's own documentation says a reservation does not reserve capacity. Deploy first, measure the load, then reserve.
  • "Unified already covers Azure OpenAI." Ask for the share of the Unified fee driven by Azure consumption and compare it with $12,000 a year for Professional Direct.

How does the answer change between a pilot and a production rollout?

For a pilot on Standard deployments, keep costs low and collect data. Standard support is enough, the availability SLA is the only commitment that applies, and the goal is a measured tokens per minute figure you can size a reservation on.

Production on provisioned capacity

Once a workload is customer facing or feeds an operational process, move the latency sensitive path to PTUs so the Latency Target and its metrics exist. Buy Professional Direct for the two hour Sev B. Negotiate the response time clause, the pinned target and retirement notice before the first reservation, because that is when Microsoft wants the commitment most.

What to do next

  1. Before the negotiation. Stop pushing the availability percentage. Ask for a credit on the published Latency Target, pinned to the Effective Date so a model retirement cannot reset it downward.
  2. During the pilot. Measure tokens per minute and size PTUs on that figure, and hold sustained usage above 60 percent before adding more.
  3. Before every model migration. Rerun the PTU sizing, because the new model can change the capacity you need several times over.
  4. At the support decision. Buy for the Sev B number. Professional Direct's two hours against Standard's four is the only support difference an AI incident will feel.
  5. Before a Unified renewal. Price Azure support separately against Professional Direct, since Unified charges a percentage of the fastest growing AI line.
  6. Before go live. Assign an owner for SLA claims, rehearse the claim path and make opening a support case during any outage a rule. Our Microsoft practice can run this with your team.
When to bring in help

Want a second opinion on your Microsoft licensing? Our Microsoft licensing consultants work only for buyers, with no reseller margin.

Frequently asked questions

What does the Azure OpenAI SLA cover, and what does it leave out?

Availability only. The 99.9 percent monthly uptime commitment asks whether the endpoint returned a response, and nothing more. Time to first token, throughput under burst load and output quality are outside it. The separate 99 percent latency SLA on provisioned deployments is a floor on token rate, which is a weaker thing than a guaranteed response time.

How much does an Azure OpenAI outage pay?

At most 25 percent of that month's Azure OpenAI charge, as a credit on a later invoice. If Azure OpenAI bills you $180,000 a month, a 95 minute outage returns $18,000 and an eight hour outage returns $45,000. Your other Azure services are not credited, even if the outage affected applications running on them.

How long do I have to claim an Azure OpenAI SLA credit?

Two months after the end of the billing month in which the incident occurred. Microsoft 365 and other online services close a month earlier, at the end of the following calendar month. Put both dates in the incident runbook, and log the support case number you opened during the outage, since the claim must describe your attempts to resolve it.

Does Azure OpenAI have a latency SLA or only an uptime SLA?

Yes, for provisioned deployments. Microsoft announced a 99 percent token generation latency SLA in November 2024 and lists a Latency Target for each model, such as 25 tokens per second for gpt-4o and 50 for the gpt-5 family. Standard deployments have no latency commitment, so any response time promise to your users rests on provisioned capacity.

Which Azure support plan does an AI workload need?

Standard at $100 a month for pilots and internal tools, since it already answers a Sev A in under an hour. For customer facing AI, Professional Direct at $1,000 a month answers the Sev B cases that latency problems produce in two hours instead of four. Developer is only suitable for test subscriptions, as it handles Sev C cases alone.

Should Azure OpenAI support go through Microsoft Unified?

Only after you have priced the alternative. Unified fees scale with your Microsoft spend, so fast growing Azure OpenAI consumption raises them automatically. At roughly 8 to 10 percent, $2.16 million a year of AI spend adds $172,800 to $216,000, against $12,000 a year for Professional Direct. Ask Microsoft to quote Unified with and without the AI consumption.

Does an Azure OpenAI PTU reservation guarantee capacity?

No. A reservation is a billing discount, applied to PTUs you have already deployed. Capacity for a model in a region can run out, so Microsoft tells customers to create the deployment first and buy the reservation afterwards. If you buy first, you can end up paying for a reservation you cannot use.

Newsletter
Licensing news that changes what you pay

One email a week on vendor price moves, audit activity and what worked in recent renewals.

Subscribe
Vendor Shield
An advisor on call for every vendor conversation

Always on advisory for renewals, audits and contract questions across your software vendors.

Explore Vendor Shield
Advisory White Paper

Get the Azure OpenAI commitment guide.

The SLA remedies, the latency clause, PTU sizing and the support plan decision, worked through end to end.

Gated with a work email on the download page. No sales follow up you did not ask for.

Get the White Paper →
We never share your details with vendors.

Microsoft licensing news, once a week.

Price changes, audit activity and what worked in recent renewals. No vendor spin.