HomeMicrosoft HubAzure OpenAI Pricing
Microsoft  |  Azure OpenAI Pricing Buyer Guide 2026

Azure OpenAI pricing, the model tier is the number nobody negotiates

Two meters produce the bill: token meters by model on Standard, and provisioned throughput units by the hour. A reservation reprices the second meter, and nothing reprices the first except your model choice and a negotiated rate, which is why the largest single number on the page is the one nobody has to negotiate for.

Prepared by Redress Compliance · August 7, 2026 · Microsoft and GenAI advisory. Based on 20 to 30 Azure OpenAI pricing reviews run 2024 to 2025.

Executive summary

Wrong model, not wrong rate, was the biggest finding. In about half the estates, a single workload parked on a frontier tier where a mini tier passed acceptance testing cost more than every negotiated discount added together.

The worked routing example: a $32,850 monthly workload on GPT-4o dropped to $9,460.80, a 71.2 percent reduction, by splitting traffic between GPT-4.1 and GPT-4.1 mini, before anyone from Microsoft was in the room.

Recovered spend across the reviews ran 25 to 40 percent of the pre optimization run rate.

The PTU breakeven is a fixed number, and the monthly term cannot pay.

Microsoft sets each model's PTU output weighting equal to its Standard price ratio, so Standard costs the input rate per normalized million tokens whatever the prompt mix.

Making the breakeven computable before the first meeting: the one month reservation breaks even at 95.0 to 107.0 percent of the capacity bought, and the one year at 80.7 to 90.9 percent sustained.

PTU at list buys a latency floor and reserved capacity, not cheaper tokens, and reservations sized on pilot peaks measured 20 to 40 percent utilization by month three. At that level the reservation is a donation.

The yearly discount is 15 percent, not the 30 the account team quotes. Twelve monthly reservations cost 12 times $260, which is $3,120 per PTU, against $2,652 on the one year term: exactly 15 percent cheaper.

And prompt caching is worth 90 percent on the GPT-5 tiers and 75 percent on GPT-4.1 and o4-mini; the 50 percent figure still in circulation is the GPT-4o number, two generations stale.

The MACC makes the reservation a renewal instrument.

Azure OpenAI is a first party service, so tokens and PTU reservations both draw down the Microsoft Azure Consumption Commitment, and buyers who lined the one year PTU term up with the MACC anniversary got drawdown treatment nobody had offered them.

The commitment conversation belongs in the EA cycle, not the AI project budget.

71.2%
The worked routing reduction, $32,850 to $9,460.80 monthly, before any negotiation.
95 to 107%
The one month PTU breakeven against capacity bought: it cannot pay at list.
80.7 to 90.9%
The sustained utilization the one year PTU reservation needs to beat Standard tokens.
25 to 40%
Recovered spend against the pre optimization run rate across our pricing reviews.
1.

The Global Standard list, July 2026, and the four levers

ModelInput per 1MOutput per 1MCaching discount
GPT-5.4$2.50$15.0090 percent
GPT-5$1.25$10.0090 percent
GPT-5 mini$0.25$2.0090 percent
GPT-5 nano$0.05$0.4090 percent
GPT-4.1$2.00$8.0075 percent
GPT-4.1 mini$0.40$1.6075 percent
o4-mini$1.10$4.4075 percent
GPT-4o$2.50$10.0050 percent

Four levers move the token bill, largest first. Model tier, where the input spread between GPT-4o and GPT-4.1 mini runs $2.50 against $0.40, routed per workload rather than per estate. Prompt caching, at 10 percent of the input rate on GPT-5 tiers for repeated prefixes on high frequency endpoints.

Batch, deployed to GlobalBatch for 50 percent off against a 24 hour window. And context discipline, because on an 8 to 1 input heavy workload the input meter is most of the bill. Data Zone Standard lists exactly 10 percent above Global on every row.

2.

The PTU math, published rates and real breakevens

Microsoft publishes the provisioned rates, $1.00 per PTU hour on Global Provisioned, $260 per PTU on the one month reservation, $2,652 on the one year, and most buyers never see them because the account team arrives with a bundled annual number instead of a unit price.

The breakeven arithmetic is mechanical: because the PTU output weighting matches the Standard price ratio, the one month term breaks even at 95 to 107 percent of purchased capacity, impossible in practice, and the one year at 80.7 to 90.9 percent sustained.

Achievable only by flat production serving.

What PTU legitimately buys is the latency floor and guaranteed capacity in constrained regions, priced as insurance rather than savings, and sized to measured sustained load, never the pilot peak that measured 20 to 40 percent utilization by month three.

The commitment stack above it, where the reservation meets the MACC and the EA, is worked in the Azure agreement negotiation.

Free white paper

The enterprise AI contract negotiation playbook

The AI meter clause set across the estate: the reservation math, the commitment drawdowns, the caching and routing disciplines, and the terms that survive model generations.

Get the white paper →
3.

The routing discipline, worked end to end

The worked example transfers to any estate: 8,760 million input and 1,095 million output tokens monthly costs $32,850 on GPT-4o; moved wholesale to GPT-4.1 it costs $26,280, exactly 20 percent off for a model string change.

And split, with the 80 percent of traffic that is classification, extraction, and routing on GPT-4.1 mini and the remainder on GPT-4.1, it costs $9,460.80.

The only inputs are your token volumes and two rates from the table, which is why the routing audit runs before any negotiation: the case study estate moved an $80,000 monthly bill to $55,000, 31 percent, with the model map doing most of the work.

The comparison shopping runs the same way, the Bedrock pricing guide and the OpenAI direct playbook pricing the same models under different wrappers.

Try Vera AI · free 30 day trial
Vera maps your workloads to the cheapest passing tier in minutes.
  • Percentile standing for your exact deal size and industry, from real closed transactions
  • Scenario simulation before the call: test alternative terms and see the financial impact of each
  • A negotiation playbook, talking points, and a two page executive brief on day one
Start the free Vera AI trial →30 days free · no credit card · cancel anytime
4.

What we saw across pricing reviews, 2024 to 2025

Across roughly 20 to 30 Azure OpenAI pricing reviews Fredrik Filipsson ran between 2024 and 2025, recovered spend landed between 25 and 40 percent of the pre optimization run rate, ranked by value:

Half
Estates with a misrouted workload

A frontier tier where a mini tier passed acceptance, worth more than every discount combined.

20 to 40%
PTU utilization by month three

On reservations sized to pilot peaks: at that level the reservation is a donation.

The MACC finding rounds out the file: buyers who timed the one year PTU term to the commitment anniversary got drawdown treatment nobody had volunteered, because Azure OpenAI consumption counts toward the MACC and the reservation is therefore a renewal instrument wearing an infrastructure label.

The governance layer that keeps the routing map current, budgets, alerts, and the monthly model mix review, sits in the Azure FinOps framework, and the operational floor underneath, quotas and regional capacity, in the SLA and support analysis.

5.

Your first five moves

  1. Run the routing audit first: every workload against the cheapest tier that passes acceptance, where the 25 to 40 percent lives.
  2. Compute the PTU breakeven before the meeting, because it is a fixed number and the monthly term cannot clear it.
  3. Size any reservation to measured sustained load, never the pilot peak, and take the one year term only above the 81 to 91 percent line.
  4. Exploit the current caching rates, 90 percent on GPT-5 tiers, and deploy asynchronous work to Batch at half price.
  5. Line the reservation up with the MACC anniversary and negotiate it inside the EA cycle. The Microsoft practice runs the math with you.
6.

Frequently asked questions

How is Azure OpenAI priced?

Two meters: token consumption by model on Standard deployments, GPT-5 at $1.25 input and $10 output per million tokens down to GPT-5 nano at $0.05 and $0.40, and provisioned throughput units at $1.00 per hour, $260 per PTU monthly reserved, or $2,652 yearly.

A reservation reprices the second meter; only model choice and negotiated rates reprice the first.

Do Azure OpenAI PTU reservations save money?

Rarely at list: the one month reservation breaks even at 95 to 107 percent of purchased capacity, which cannot happen, and the one year at 80.7 to 90.9 percent sustained utilization, achievable only by flat production serving.

PTU buys a latency floor and guaranteed capacity, priced as insurance, and reservations sized on pilot peaks measured 20 to 40 percent utilization by month three.

What is the biggest Azure OpenAI cost saving?

Model routing: in half our estates a single workload on a frontier tier, where a mini tier passed acceptance testing, cost more than every negotiated discount combined.

The worked example cut $32,850 monthly to $9,460.80, 71.2 percent, by splitting traffic between GPT-4.1 and GPT-4.1 mini, and recovered spend across reviews ran 25 to 40 percent.

How much does Azure OpenAI prompt caching save?

90 percent on the GPT-5 tiers, 75 percent on GPT-4.1 and o4-mini, and 50 percent on GPT-4o, the stale figure still circulating.

Cached prefixes expire after minutes of inactivity, so caching pays on high frequency endpoints and does nothing for overnight runs; Batch covers those instead at 50 percent off against a 24 hour window.

Does Azure OpenAI count toward a MACC?

Yes: it is a first party Azure service, so token consumption and PTU reservations both draw down the Microsoft Azure Consumption Commitment.

That makes the one year PTU term a renewal instrument, and buyers who aligned it with the MACC anniversary got drawdown treatment nobody had offered, which is why the reservation negotiates inside the EA cycle.

Is the yearly PTU reservation 30 percent cheaper?

No: twelve monthly reservations cost $3,120 per PTU against $2,652 on the one year term, exactly 15 percent, whatever the account team quotes.

The deeper arithmetic matters more: the yearly term only beats Standard tokens above 80.7 to 90.9 percent sustained utilization, so the discount question is secondary to whether the reservation should exist at all.

© 2026 Redress Compliance · Independent, buyer sideredresscompliance.com
Industry Recognized
500+ Enterprise Clients
$2B+ Under Advisory
11 Vendor Practices
100% Buyer Side Independent
Enterprise AI White Paper

The full enterprise AI contract negotiation playbook from the AI practice.

The AI meter clause set: the reservation math, the commitment drawdowns, the caching and routing disciplines, and the terms that survive model generations.

Gated with a work email on the download page. No sales follow up you did not ask for.

Get the White Paper →
Independent, buyer side. We never share your details with vendors.
Run the software spend health check against your Azure estate in under five minutes.
Open the Tool → Microsoft Advisory →
Editorial boardroom interior

The advisor your vendors do not want.

500+ enterprise clients. 11 vendor practices. Industry recognized. One conversation can change what you pay for the next three years.

Stay ahead of Microsoft pricing and contract moves.

One buyer side briefing a week. Renewal signals, discount bands, and the levers that work. No vendor spin.