Data center aisle lined with server racks
Azure OpenAI

Azure OpenAI pricing in 2026. Why the model tier matters more than the discount.

Token rates by model, what provisioned throughput units cost, when a reservation beats Standard pricing, and how Azure OpenAI spend fits into your MACC and EA renewal.

Contact Us Microsoft Advisory
500+Enterprise clients
$2B+Under advisory
PublishedJuly 26, 2026UpdatedSeptember 23, 2026
ContentsKey takeawaysHow Azure OpenAI is pricedSavings from model routingWhat PTUs costPTU breakeven against StandardAzure OpenAI and your MACCAzure OpenAI or OpenAI directWhat our reviews foundWhat to do nextFAQ

Azure OpenAI bills tokens by model on Standard and PTUs by the hour or by reservation. At list prices a PTU reservation rarely beats Standard, so model routing, caching and Batch cut the bill before any negotiation starts.

Key takeaways
  • Two meters. Standard deployments bill tokens at a rate set by the model; provisioned deployments bill PTUs at $1.00 an hour, $260 a month reserved or $2,652 a year reserved.
  • Routing saves the most. Splitting one worked workload between GPT-4.1 and GPT-4.1 mini cut it from $32,850 to $9,460.80 a month before any negotiation.
  • Reservations rarely pay at list. A PTU is priced close to the Standard token value of the throughput it delivers, so a reservation wins only under near constant, full load.
  • Deploy first, reserve second. A PTU reservation is a billing discount that guarantees no regional capacity, so confirm capacity with a live deployment before you buy.
  • Caching pays more on newer models. Cached input discounts are larger on the GPT-5 tiers than on GPT-4.1 or GPT-4o, and cost models built on GPT-4o figures understate them.
  • It draws down your MACC. Tokens and PTU reservations both count toward the commitment, so time the one year term to the MACC anniversary.
  • Recovered spend is real. Across our 2024 and 2025 reviews, buyers recovered 25 to 40 percent of their pre optimization run rate.

How is Azure OpenAI priced in 2026?

Azure OpenAI bills on two meters. Standard deployments charge per million input and output tokens at a rate set by the model you call. Provisioned deployments charge for provisioned throughput units (PTUs) by the hour, or at a discount if you reserve them for a month or a year.

A reservation reprices only the second meter. The token meter changes with the model you choose and any rate you negotiate. Most enterprise spend runs on Global Standard, and the table shows its July 2026 list.

Azure OpenAI Global Standard list prices, July 2026, per 1 million tokens
ModelInputOutputDiscount on cached input
GPT-5.4$2.50$15.0090 percent
GPT-5$1.25$10.0090 percent
GPT-5 mini$0.25$2.0090 percent
GPT-5 nano$0.05$0.4090 percent
GPT-4.1$2.00$8.0075 percent
GPT-4.1 mini$0.40$1.6075 percent
o4-mini$1.10$4.4075 percent
GPT-4o$2.50$10.0050 percent

Microsoft has since added GPT-5.5, the GPT-5.6 series and a first GPT-6 model. Check their rates and PTU figures on Microsoft's pricing and sizing pages before you extend this page's arithmetic to them.

Two other deployment types change the rate on every row. Data Zone Standard, which keeps processing inside the US or EU, lists exactly 10 percent above Global. Batch jobs deployed to GlobalBatch cost 50 percent less than Global Standard, in exchange for results within a 24 hour window.

What changes the token bill the most?

Four things set the Standard bill, in order of the money involved:

  1. Model tier. Input on GPT-4o costs $2.50 per million tokens and on GPT-4.1 mini $0.40. Pick the tier per workload, since a summarizer and a contract reviewer rarely need the same model.
  2. Prompt caching. On the GPT-5 tiers a cached prefix bills at 10 percent of the input rate. The 50 percent figure still quoted in many cost models is the GPT-4o number, two generations old. GPT-4.1 and o4-mini sit at 75 percent.
  3. Batch. Anything that can wait up to 24 hours, such as nightly document classification or backfills, costs half on GlobalBatch.
  4. Context discipline. Most enterprise workloads send far more than they receive. At an 8 to 1 input to output ratio, the input meter is most of the bill, so trimming retrieved context and repeated instructions pays directly.

Caching has conditions. A prompt needs at least 1,024 tokens, and the first 1,024 must be identical to an earlier request for a cache hit. The default cache is typically cleared after 5 to 10 minutes of inactivity, so it pays most on high frequency endpoints with a long, stable system prompt.

Microsoft also documents extended cache retention of up to 24 hours on GPT-4.1 and the full size GPT-5 models from GPT-5 to GPT-5.5, so a daily job that reuses long instructions can earn cache hits too. On GPT-5.6, the usage object reports cache writes separately, so check them before you count the saving.

Watch the briefingResearch briefing · 4:33

How to Negotiate with OpenAI and Anthropic: The Vendors With Nobody to Call

How much can model routing save on Azure OpenAI?

Routing saves more than any discount we have seen negotiated. It means sending each type of request to the cheapest model that passes your acceptance tests. The calculation needs only your token volumes and the rates in the table above.

Worked example: one workload, three model choices

Take a workload that sends 8,760 million input and 1,095 million output tokens a month, the 8 to 1 ratio above. On GPT-4o it costs $32,850 a month. Moved wholesale to GPT-4.1, a change to the model string, it costs $26,280, exactly 20 percent less.

Now split it. About 80 percent of the traffic is classification, extraction and routing, which GPT-4.1 mini handles. The remainder stays on GPT-4.1.

Monthly cost of 8,760 million input and 1,095 million output tokens
OptionInput costOutput costMonthly total
All on GPT-4o8,760 × $2.50 = $21,9001,095 × $10.00 = $10,950$32,850
All on GPT-4.18,760 × $2.00 = $17,5201,095 × $8.00 = $8,760$26,280
80 percent on GPT-4.1 mini7,008 × $0.40 = $2,803.20876 × $1.60 = $1,401.60$4,204.80
20 percent on GPT-4.11,752 × $2.00 = $3,504219 × $8.00 = $1,752$5,256
Split total$9,460.80

The split costs $9,460.80 a month, a 71.2 percent reduction from GPT-4o, before anyone from Microsoft is in the room. In one of our case studies, a customer took an $80,000 monthly bill to $55,000, a 31 percent cut, and the model map did most of that work.

How do you find misrouted workloads in your own tenant?

Azure already records what you need:

  • Azure Monitor metrics. Processed Prompt Tokens and Generated Completion Tokens, split by the ModelDeploymentName dimension, give volume and input to output ratio per deployment.
  • Prompt Token Cache Match Rate. A low rate on an endpoint with a long fixed system prompt usually means the prompt changes in its first 1,024 tokens, often a timestamp or user name placed too early.
  • Cost Management. Group Azure OpenAI cost by meter and resource to see which deployments carry the spend.
  • The usage object in each API response. It reports prompt tokens, completion tokens and cached tokens per call, so application teams can attribute cost to a feature.

Then test each high volume deployment against the next tier down on a fixed acceptance set. Start with the deployments that carry the most input tokens, since a pass there on a mini tier cuts the largest line on the bill.

Free white paper

Enterprise AI Contract Negotiation Guide

Reservation arithmetic, MACC drawdown and the contract terms to ask for on AI consumption.

Get the white paper →

What do Azure OpenAI provisioned throughput units cost?

Microsoft publishes the rates. Global Provisioned costs $1.00 per PTU per hour on hourly billing, $260 per PTU on a one month reservation and $2,652 per PTU on a one year reservation. Most buyers never see those unit prices, because the account team arrives with a bundled annual number instead.

At 730 hours a month, hourly billing comes to $730 per PTU. The monthly reservation is about 64 percent cheaper than that.

How much cheaper is the one year reservation than the monthly one?

The one year term is exactly 15 percent cheaper. Twelve monthly reservations cost 12 × $260, or $3,120 per PTU, and the one year term costs $2,652. Account teams often quote 30 percent. Whatever figure you hear, ask for the two per PTU prices and do the division yourself.

What does a PTU reservation not do?

The reservation is a billing discount on deployments, with limits Microsoft documents:

  • It does not guarantee capacity. Create the deployment first to confirm capacity exists in the region, then buy the reservation.
  • It must match the deployment type. A Global reservation does not cover a Data Zone or Regional deployment. You can exchange a reservation to another deployment type, region or term, but the term restarts on the new reservation.
  • Deployments cannot be paused. Billing stops only when the deployment is deleted. Deployed PTUs beyond the reserved quantity bill at the hourly rate.
  • Minimums apply. Every provisioned deployment has a minimum PTU count per model, and Regional minimums are higher than Global and Data Zone ones.

The reservation is model independent, so it keeps covering your deployments when you move from one model generation to the next.

When does a PTU reservation beat Standard token pricing?

At list prices, the one month reservation breaks even only at 95 to 107 percent of the capacity you bought, which no real workload sustains. The one year reservation needs 80.7 to 90.9 percent sustained load, which only flat, round the clock production serving reaches.

Why is the breakeven a fixed number?

Microsoft measures PTU capacity in input tokens per minute and counts each output token as several input tokens. That output weighting equals each model's Standard price ratio: 4 for GPT-4o and GPT-4.1, 8 for GPT-5 and 6 for GPT-5.4.

As a result, Standard costs the input rate per normalized million tokens, whatever mix of prompt and completion you send. That makes the comparison with a PTU a single division.

PTU reservation breakeven against Standard, at list prices (730 hour month)
ModelInput tokens per minute per PTUStandard value of one fully used PTU per monthOne month breakevenOne year breakeven
GPT-4o2,500$273.7595.0 percent80.7 percent
GPT-4.13,000$262.8098.9 percent84.1 percent
GPT-5.42,400$262.8098.9 percent84.1 percent
GPT-4.1 mini14,900$261.0599.6 percent84.7 percent
GPT-54,750$260.06100.0 percent85.0 percent

The first row works like this. One GPT-4o PTU handles 2,500 input tokens per minute, and over 43,800 minutes in a month that is 109.5 million normalized tokens, worth $273.75 at the Standard input rate. The monthly reservation at $260 is 95.0 percent of that, and the yearly term at $221 a month is 80.7 percent.

On hourly billing, a fully loaded GPT-4o PTU costs about 2.7 times the same tokens on Standard, so hourly PTUs never win on cost.

What is PTU actually worth paying for?

PTU buys a latency floor and reserved capacity, which matters in regions where Standard quota runs short at peak. Treat it as insurance and price it that way. Size any reservation to measured sustained load, never to a pilot peak.

Two features help with sizing. Spillover sends requests that exceed your PTUs to a Standard deployment in the same resource, so you can reserve for the base load and pay tokens for the bursts. Cached tokens do not consume PTU capacity, so with a high cache hit rate a smaller reservation carries the same traffic.

Developer working with monitoring dashboards on several screens
The V2 provisioned managed metric in Azure Monitor shows PTU load per deployment. Read it at hourly granularity over at least a month before sizing a reservation, so you can count the hours each day that load sits below the level you plan to reserve.

Why we would not buy PTUs to lower the token bill

The usual advice is to move steady workloads onto a PTU reservation to save money. We disagree, because the breakeven table leaves almost no room for savings at list prices. The pilot sized reservations we reviewed were running far below that line by month three, and every idle PTU hour is paid for in full.

Say a team reserves 100 GPT-4o PTUs on the one year term: $265,200 a year. At 30 percent sustained use, the tokens served would have cost about $98,550 on Standard, so $166,650 buys nothing but headroom. Run routing, caching and Batch first, then reserve only what latency or capacity requires.

Does Azure OpenAI spend count toward your MACC?

Yes. Azure OpenAI is a first party Azure service, so both token consumption and PTU reservations draw down the Microsoft Azure Consumption Commitment. That turns the one year reservation into part of your renewal, and the negotiation belongs in the Enterprise Agreement cycle rather than the AI project budget.

The commitment stack above the reservation, where it meets the MACC and the EA, is worked through in our Azure agreement negotiation guide, and commitment sizing in the MACC negotiation guide.

Contract terms to ask for

  • Written MACC eligibility. Confirmation that Azure OpenAI tokens and provisioned reservations count toward the commitment at full value, so a later program change cannot reclassify them.
  • Term alignment. A one year reservation start date that matches the MACC anniversary, so the next reservation decision falls inside the renewal window.
  • Price protection by model. Current Standard rates held for the models you run in production for the agreement term.
  • Capacity before commitment. Confirmed capacity in the named region before any reservation is signed, since the reservation itself guarantees none.
  • Exchange without a term reset. The standard exchange restarts the term, and self service cancellations are capped at $50,000 of commitment in any rolling 12 months. If data residency could push you from Global to Data Zone, ask for an exchange that keeps the original end date.

What will the Microsoft account team say?

  • "The annual commitment saves you 30 percent." Ask for the per PTU monthly and yearly prices in writing, then divide the yearly price by twelve monthly terms.
  • "You need PTUs for production." You need them for a latency guarantee or scarce regional capacity. Show your measured load and the spillover plan.
  • "Buy the reservation now to secure capacity." Microsoft's own documentation says reservations do not guarantee capacity. Deploy first, then reserve.
  • "This sits outside the EA discussion." It draws down the MACC, so it sits inside it.

Is Azure OpenAI cheaper than buying from OpenAI directly?

List token prices for the same models are close enough that the wrapper matters more than the rate. Through Azure, spend draws down a MACC you may already hold and sits under your existing Microsoft terms. Direct, you negotiate with a much smaller sales organization.

Buying direct from OpenAI or Anthropic

Each vendor has fewer than 50 sales reps globally, focused on $100M+ deals. Below $10M a negotiation rarely starts, and discounts run 5 to 25 percent depending on commitment size. The pressure that works is credible competition between OpenAI, Anthropic and Gemini, backed by benchmarked prices.

The OpenAI direct procurement guide and the Bedrock pricing guide price the same kind of models under different wrappers, and the Azure OpenAI versus direct comparison sets them side by side.

What have we seen in recent Azure OpenAI pricing reviews?

I ran roughly 20 to 30 Azure OpenAI pricing reviews between 2024 and 2025. Recovered spend landed between 25 and 40 percent of the pre optimization run rate. The findings, ranked by value:

  1. A misrouted workload. In about half the customers, one workload ran on a frontier tier where a mini tier had passed acceptance testing. That single workload cost more than every negotiated discount combined.
  2. Oversized reservations. Reservations sized to pilot peaks were running at 20 to 40 percent use by month three.
  3. Unused MACC timing. Buyers who lined the one year PTU term up with the MACC anniversary got drawdown treatment Microsoft had not offered them unprompted. Those who bought on the project's schedule did not.
The largest number on an Azure OpenAI bill is the model tier, and changing it needs no one's signature at Microsoft.

Budgets, alerts and the monthly model mix review that keep the routing map current are covered in the Azure FinOps cost governance guide. Quotas and regional capacity, the operational floor under all of this, are covered in the Azure OpenAI SLA and support analysis.

What to do next

  1. Run the routing audit first. Test every high volume workload against the cheapest tier that passes acceptance. That is where most of the recoverable spend sits.
  2. Compute the PTU breakeven before the meeting. Use the tokens per minute figure for your model and the Standard input rate. The monthly term cannot clear it at list.
  3. Size any reservation to measured sustained load. Take the one year term only when sustained use sits above the roughly 81 to 91 percent line for your model, and use spillover for bursts.
  4. Use the current caching and Batch rates. Put stable instructions at the front of the prompt to earn the GPT-5 cache discount, and move anything that can wait 24 hours to Batch at half price.
  5. Deploy first, reserve second. Confirm capacity in the region with a live deployment before you buy.
  6. Line the reservation up with the MACC anniversary. Negotiate it inside the EA cycle. Our Microsoft practice can run these numbers with you.
When to bring in help

Is a Microsoft renewal or new agreement coming up? Our Microsoft EA negotiation team works only for buyers, for a fixed fee or 25 percent of what we save you.

Frequently asked questions

How is Azure OpenAI priced?

Standard deployments charge per million input and output tokens at a rate set by the model, from GPT-5 at $1.25 input and $10 output down to GPT-5 nano at $0.05 and $0.40. Provisioned deployments charge per PTU at $1.00 an hour, $260 a month reserved or $2,652 a year reserved.

Do Azure OpenAI PTU reservations save money?

Rarely at list prices. The monthly term breaks even at 95 to 107 percent of purchased capacity and the yearly term at 80.7 to 90.9 percent sustained load. Reservations sized on pilot peaks that we reviewed measured 20 to 40 percent use by month three, so buy PTUs for latency and capacity, not savings.

What is the biggest Azure OpenAI cost saving?

Model routing. In about half the customers we reviewed, a single workload running on a frontier tier where a mini tier passed acceptance testing cost more than every negotiated discount combined. Fixing it takes acceptance test results and a change to the model string.

How much does Azure OpenAI prompt caching save?

Cached input costs 90 percent less on the GPT-5 tiers, 75 percent less on GPT-4.1 and o4-mini, and 50 percent less on GPT-4o. The discount applies only to the repeated prefix, never to output tokens, so a chat workload with short prompts and long answers gains little. Batch at 50 percent off within 24 hours is often the bigger saving there.

Does Azure OpenAI count toward a MACC?

Yes. As a first party Azure service, both token consumption and PTU reservations draw down the Microsoft Azure Consumption Commitment. Get that confirmed in writing before you sign, and negotiate the reservation as part of the EA cycle so its term ends inside your next renewal window.

Is the yearly PTU reservation 30 percent cheaper?

No. The one year reservation is exactly 15 percent cheaper than twelve monthly reservations, whatever figure the account team quotes. The bigger question is whether the reservation should exist at all, since the yearly term beats Standard tokens only when sustained load stays above the breakeven line for your model.

What is the minimum PTU purchase on Azure OpenAI?

Global Provisioned and Data Zone Provisioned deployments start at 15 PTUs and grow in steps of 5. At list, that entry point costs $3,900 a month on the monthly reservation or $39,780 a year on the one year term. Regional Provisioned starts at 25 or 50 PTUs depending on the model.

Newsletter
Licensing news that changes what you pay

One email a week on vendor price moves, audit activity and what worked in recent renewals.

Subscribe
Vendor Shield
An advisor on call for every vendor conversation

Always on advisory for renewals, audits and contract questions across your software vendors.

Explore Vendor Shield
Advisory White Paper

Get the enterprise AI contract negotiation guide.

The clauses that matter for AI spend: reservation arithmetic, commitment drawdown, caching and routing controls, and terms that survive a change of model generation.

Gated with a work email on the download page. No sales follow up you did not ask for.

Get the White Paper →
We never share your details with vendors.

Microsoft licensing news, once a week.

Price changes, audit activity and what worked in recent renewals. No vendor spin.