Project team walking through a plan in a boardroom session
Google Cloud · GPU and TPU Capacity · Negotiation Playbook

Negotiating Google Cloud When GPU and TPU Capacity Is the Scarce Good

In a supply-constrained accelerator market, a 55% discount on chips you cannot get delivered is worth nothing, and Google's own earnings calls give you the proof that it is rationing rather than selling. This page shows how to convert commitment length into written allocation, named regions, and dated delivery with remedies, and what the numbers look like when you get it right.

Contact Us Google Cloud Advisory
500+Enterprise clients
$2B+Under advisory
Industry Recognized
500+ Enterprise Clients
$2B+ Under Advisory
11 Vendor Practices
100% Buyer Side Independent

In a supply-constrained accelerator market, a 55% discount on chips you cannot get delivered is worth nothing, and Google's own earnings calls give you the proof that it is rationing rather than selling. This page shows how to convert commitment length into written allocation, named regions, and dated delivery with remedies, and what the numbers look like when you get it right.

Google Has Already Told the Market It Cannot Serve Everyone

Your opening position was written for you by Alphabet's investor relations team. CFO Anat Ashkenazi told analysts on the July 22, 2026 call that Google remains in a "supply-constrained environment," and made a point of noting the company has said so for multiple quarters running. Cloud backlog stands at $514 billion, up more than $50 billion sequentially and up from $106 billion four quarters earlier, with over half expected to convert to revenue inside 24 months. Read that as a buyer, not an analyst: Google has already sold roughly a quarter of a trillion dollars of capacity it has to physically deliver in the next eight quarters, and it is renting third-party infrastructure as what Ashkenazi called a "bridging strategy" to cover the gap. Then Pichai supplied the detail that actually decides your outcome: baseline capacity is allocated first to frontier AGI development. That is the internal queue, in the CEO's own words, and unless your allocation is written into a contract with a region and a date, you are behind DeepMind and Search in it. The tactical consequence is that your rep cannot hold two positions at once. If capacity is freely available, then a written allocation clause with a delivery date costs Google nothing and should be signed. If it is not, then the rep has just conceded that unit price is the wrong thing to be arguing about. Force the choice early, in writing, and get it into the record before you discuss rates at all. Through 2026 and 2027, the contested asset is the delivery date.

A 55% discount on chips that arrive in the third quarter of next year is not a discount, it is a queue ticket.

Read the Rate Card as a Map of What Google Wants From You

The published TPU rate card is not a price list, it is a statement of intent. Ironwood in us-central1 sits at $12.00 on-demand per chip-hour, $8.40 for a one-year commitment, and $5.40 for three years. DWS Flex-start is $6.00. Sit with that for a moment: the three-year commitment prices below spot-style flex capacity. Google is not paying you to consume more, it is paying you to consume for longer, and it is willing to sell duration below the price of uncertainty to get it. Trillium gives the cleanest public benchmark of the same behavior at $2.70 on-demand, $1.89 for one year, and $1.22 for three, which is roughly 30% for one year and 55% for three. The financial logic behind it is on the same earnings call: capex guidance raised three consecutive quarters to $195 to $205 billion, and free cash flow at negative $5.9 billion against $39.1 billion of operating cash flow. Google is funding a construction program against contracted revenue, so what the sales motion is actually compensated for is term length and revenue certainty, not incremental volume. That tells you exactly what currency you hold and what to buy with it. Do not spend three-year duration on another five points of discount, because the discount curve is published and your rep will concede it without a fight. Spend it on named regions, chip-count floors, and dated delivery with remedies. Sequencing matters here, and the mechanics of when to apply which lever are covered in our guide on when to start a Google Cloud negotiation and how to sequence pressure.

Consumption mode, Ironwood us-central1 Price per chip-hour What Google is signaling
On-demand$12.00List anchor, not a real target
DWS Flex-start$6.00Queued capacity, no assurance
DWS Calendar mode$8.40The published price of certainty
1-year commitment$8.40Duration matched to calendar assurance
3-year commitment$5.40Priced under flex to buy your term

Note the London comparison as well: Ironwood at $13.20 on-demand, $9.24 for one year, and $5.94 for three. Region is a live price variable, so treat any offer to relocate your allocation as a rate change and price it accordingly. The broader accelerator levers, including how these commitments interact with Vertex and model consumption, are set out in our Google Cloud AI contract playbook.

The Price of Certainty Is Already Published, So Use It

When a rep tells you a written allocation clause is "not something we do," point at Google's own Dynamic Workload Scheduler rate card. Google already sells certainty as a line item. Calendar mode reservations, which hold the machines for you on dates you name, list at $90.22 per hour for a4-highgpu-8g against $64.44 for Flex-start, where you queue and hope. That is a 40% premium for the right to know when you get the hardware. The a3-ultragpu-8g pair runs $59.36 versus $42.40, another 40%. The a3-megagpu-8g pair narrows to $44.00 versus $40.32, roughly 9%, which tells you the premium tracks how scarce the silicon actually is: newest generation, highest certainty tax. Note also that calendar mode only exists for the 8-GPU shapes. If your workload wants a fractional shape, Google's own product line does not offer you a way to buy assurance at any price, which is precisely why the assurance has to come from contract language instead.

Shape Calendar (assured) Flex-start (queued) Certainty premium
a4-highgpu-8g$90.22/hr$64.44/hr~40%
a3-ultragpu-8g$59.36/hr$42.40/hr~40%
a3-megagpu-8g$44.00/hr$40.32/hr~9%
a3-highgpu-8g$41.60/hrnot listedn/a

Use the delta as an argument, not as a purchase. A three-year commitment already pays Google the premium: on Trillium, one year takes $2.70 to $1.89 and three years to $1.22, roughly 55% off, and Ironwood's three-year rate of $5.40 in us-central1 sits below the $6.00 Flex-start price. You have prepaid the certainty. The correct position is that allocation language follows the commitment into the contract at commit rates, with no calendar uplift stacked on top. Sequencing matters here; see when to start and how to stage the pressure.

Google publishes a 40% premium for assured GPU capacity, then argues assurance cannot be written into a contract.

Region Substitution Is a Price Increase Wearing a Logistics Costume

"We can place you in europe-west2 instead" is not a scheduling accommodation. On Ironwood, that move takes on-demand from $12.00 to $13.20 per chip-hour and the one-year rate from $8.40 to $9.24. On Trillium, us-east1 and us-east5 run $2.70, europe-west4 runs $2.97, and asia-northeast1 runs $3.24: roughly 10% and 20% uplifts respectively. A rep offering region flexibility is offering you a 10% to 20% rate increase and describing it as help. Price it as a rate change every time, because that is what it is.

The second-order costs are usually larger than the headline uplift and buyers routinely forget to model them. Cross-region egress on training data sets moves real money when a corpus is measured in petabytes and lives where the substitute region is not. Inference latency to endpoints your product already promises can break a customer SLA you signed independently. Data residency commitments made to a regulator or to your own enterprise customers do not bend because Google's Iowa capacity is full. And there is a fourth question almost nobody asks: Google's CFO said the company will expand third-party capacity use as a bridging strategy, so ask in writing whether your committed allocation may be fulfilled on rented infrastructure, and what that does to residency, egress path, and support response.

The ask is narrow and testable. Name the regions in the order form itself, not in an architecture appendix. Add written confirmation that any substitution requires your prior written consent, cannot carry a rate uplift, and that if Google substitutes a higher-priced region you pay the originally named region's rate. Google will counter with a "region flexibility for faster delivery" trade. Take it only if the flex regions are enumerated and rate-locked to the cheapest named region. This belongs in the same review pass as your other Google Cloud contract terms, because a substitution right without a rate lock quietly repeals whatever discount you just won.

Trading Commitment Length for Written Allocation, Not for Discount

The standard Google play is to answer a three-year commitment request with a discount curve. Refuse the trade on those terms. The published curve already gives you 55% at three years on Ironwood (us-central1 moves from $12.00 on-demand to $5.40 committed), so duration buys nothing you cannot get from the rate card. What duration should buy is a schedule. Ask for six specific things in the agreement body, not in a slide: the chip family and generation named (Ironwood or Trillium, not "TPU capacity"), a chip-count floor by calendar quarter, a named region and zone set with a defined substitution process, a first-delivery date, a ramp schedule tied to that date, and a priced remedy if the ramp slips. The zone specificity matters because allocation is granted at zone level and because Google has said publicly it will bridge shortfalls with third-party capacity, which is a question you want answered in writing rather than discovered during a data residency review.

What you concede is real, so concede it deliberately. A 36-month term, a spend floor with quarterly minimums rather than an annual lump, a reference or design-partner role for the newer generation, and a willingness to accept a mixed Trillium plus Ironwood profile are all cheap for you and valuable to a seller whose free cash flow went negative while capex guidance rose three quarters running. What is not cheap is "commercially reasonable efforts." That phrase converts your commitment into a firm obligation and Google's into an aspiration, and it is where most accelerator deals quietly fail. Price the slip instead. In our experience across recent accelerator negotiations, three remedies are achievable: service credits scaled to the undelivered chip-hours, commit relief that reduces the quarterly minimum by the shortfall percentage, and a conversion right that turns unfulfilled allocation into flexible spend usable on any GCP service rather than expiring. The Dynamic Workload Scheduler rate card gives you the arithmetic to defend this: calendar mode on a4-highgpu-8g runs $90.22 per hour against $64.44 for flex-start, roughly a 40% premium purely for certainty of timing. Google has published the price of certainty. Make it pay for the remedy. Sequence the ask against the vendor's own quarter using the guidance in when to start a Google Cloud negotiation and how to sequence pressure.

Google has already published the price of certainty at a 40% premium, so it cannot argue that a written delivery date costs nothing to promise.

What a Strong Outcome Looks Like in Numbers

Grade your own deal against public anchors before you accept a "best and final." One-year accelerator commitments should land at 30% to 37% off on-demand, three-year at 40% to 55%. Trillium's published ladder ($2.70 on-demand, $1.89 at one year, $1.22 at three) sits at the top of that band, so a three-year offer under 45% on a comparable SKU is below market. Private agreements above roughly $1 million in annual spend should reach 40% to 50% on accelerator SKUs, and the published CUD ceilings (55% general compute, 70% memory-optimized) define where the discount conversation stops and the allocation conversation starts. If you are buying at genuine hyperscale, SemiAnalysis estimates Anthropic pays about $1.60 per chip-hour for Ironwood, which is an analyst figure rather than a Google number, but it establishes that the newest generation can price below the older generation's list. Anyone quoting you list plus a haircut is not negotiating at that level.

Metric Acceptable Strong
1-year accelerator discount30%37%
3-year accelerator discount40%55%
Private agreement, $1M+ annual40%50%
Compute under commitment (portfolio)58% (median)85%+ (top quartile)
Accelerator spend under commitment30% to 40%50%+, only once allocation is written
Delivery slip remedyService creditsCredits plus commit relief plus conversion to flexible spend

Coverage deserves its own target, and it is the one place where the top-quartile number is the wrong goal. Median GCP enterprises run about 58% of compute under commitment; the top quartile exceeds 85%. For general compute, chase the top quartile. For accelerators, cap yourself lower until the allocation language is signed, because an uncovered burst hour at $12.00 is cheaper than a stranded three-year commit at $5.40 on chips that never arrived in your zone. Sizing and clause structure interact here, and the broader term set is covered in Google Cloud contract terms and how to negotiate them. Start by pulling your last four quarters of accelerator consumption by region and zone, then send Google a written allocation request with quarterly chip counts and dates before you discuss price at all.

What Google Will Do in Response, and How to Answer It

Every countermove is predictable because the rep is compensated on committed dollars, not on delivery dates, and the allocation stack sits with a capacity organization the rep does not control. Expect five plays. First, "we cannot guarantee capacity, only quota." Answer: then the commitment shrinks to match what is guaranteed. If Google will only underwrite 300 of the 1,000 Ironwood chips you asked for, the three-year commit covers 300 chips and the rest runs on-demand or Flex-start. Reps hate that trade because it cuts committed dollars by two-thirds, which is exactly why it produces an allocation schedule within two weeks. Second, a deeper unit discount instead of allocation language. A 58% discount on chips that arrive in Q3 2027 has a present value of zero, and you should say so in those words. Third, pushing you to calendar mode as the substitute for a contract clause. Price it out loud: calendar mode on a4-highgpu-8g is $90.22 per hour against $64.44 Flex-start, roughly a 40% premium for certainty. If you are also signing a three-year commit, you are paying for certainty twice, once in duration and once in the SKU. Fourth, region flexibility framed as a favor. Amsterdam is about 10% over us-east1 on Trillium and Tokyo about 20%, so any substitution right should carry price protection at the original region's rate. Fifth, a pivot to a Vertex AI or Gemini bundle to change the subject; keep those on a separate paper trail, since platform spend is where margin hides and the accelerator conversation is where your risk sits. Escalate above the rep the moment allocation is described as "not something we do," because it is a capacity decision, not a sales one. And note the asymmetry in going quiet: in a supply-constrained market, silence costs you a queue position, not a discount point, so time pressure carefully using the timing and pressure sequence rather than by stalling.

What to Do First

Spend the first week on data, not on Google. Pull twelve months of accelerator consumption by SKU, chip generation, and region, then split it into training bursts and steady-state inference, because those two profiles buy differently: bursts belong in Flex-start and calendar mode, steady inference is what you commit. Week two, write the ask as a one-page term sheet before anyone quotes a price: chip counts by generation, quarters of availability, named regions, delivery dates, and the remedy if a date slips (service credits, release from the corresponding commitment tranche, or both). Do not let the first conversation be about discount percentage; whoever sets the agenda sets the currency. Week three, send one written question and keep the reply: may committed capacity be fulfilled on third-party infrastructure, and if so, what happens to data residency, egress charges, and support response times? Google has told investors it is leaning on rented capacity as a bridge, so this is a live risk, not a hypothetical. Week four, land the allocation request ahead of quarter-end rather than inside it, because capacity planners work on a different calendar than the sales team, and the quarter-end versus year-end dynamic moves price, not supply. Two guardrails to hold all the way to signature. Cap accelerator commitment coverage at roughly 50% to 60% of forecast demand until the allocation schedule is executed, and keep a second-source option genuinely warm (a competing hyperscaler or a neocloud with dated availability), because the only real alternative to Google's queue is another queue you already stand in.

Frequently asked questions

Can Google Cloud actually contract a guaranteed number of GPUs or TPUs?

Yes, but not in the standard order form and not without a real commitment behind it. Reserved capacity and calendar mode reservations already create contractual entitlement to specific machine shapes, so the mechanism exists; the negotiation is about extending it to a named chip generation, region set, and quarterly chip-count floor with a remedy for late delivery. Expect the first draft to come back as quota language rather than allocation language, which is not the same thing.

Is a three-year TPU commitment safe when new generations launch every year?

It is safe only if the contract lets you move the commitment onto newer silicon at equivalent value. Ironwood reached general availability in April 2026 and Trillium is still on the rate card, so a three-year lock on a single generation risks paying for last year's chip at this year's price. Ask for generation portability so committed spend can be applied to whatever accelerator family is current, and treat the absence of that clause as a reason to shorten the term.

How much discount should we expect on accelerator commitments?

Public commit ladders point to roughly 30% for one year and 55% for three years on TPU chip-hours, and third-party benchmarks put negotiated private agreements at 40% to 50% or better once annual spend passes about $1 million. Published CUD ceilings run to 55% for most machine series and 70% for memory-optimized. If your offer is inside those bands but carries no allocation commitment, you have won the price argument and lost the delivery one.

Should we take Dynamic Workload Scheduler flex-start instead of committing?

Flex-start is the right instrument for genuinely bursty experimentation because it queues rather than failing on stockout, and it prices well below on-demand. It is the wrong instrument for a production training schedule with a date attached, because queueing is not a delivery guarantee. Note the arbitrage on Ironwood: the three-year commit rate sits below the flex rate in us-central1, so for predictable load the commit is cheaper and more certain at the same time.

What happens if Google fulfills our capacity on third-party infrastructure?

Alphabet has publicly described expanded third-party capacity use as a bridging strategy while it builds its own data centers, which means fulfillment location is a live question, not a hypothetical. Ask in writing whether your committed allocation can be served from non-Google facilities, and what that does to data residency commitments, egress charges, and support response obligations. If the answer is yes or unclear, get consent rights and a rate protection so a fulfillment decision cannot become a cost or compliance event on your side.

Does going quiet on the rep help in a capacity-constrained negotiation?

Less than it does in a pure price negotiation. Silence works when the vendor needs your signature to make a number; when the vendor is rationing a scarce asset, silence can cost you a slot in the allocation queue rather than winning a concession. Keep the commercial conversation slow but keep the capacity request active and dated, and escalate on delivery terms rather than on discount.

Free White Paper

What Gemini for Workspace adds to your bill

What Gemini for Workspace really costs as a Workspace add on: named user licensing, bundling pressure, and the buyer side levers that cap the spend.

Gated with a work email on the download page. No sales follow up you did not ask for.

Get the White Paper →
Independent, buyer side. We never share your details with vendors.
Negotiating Google Cloud right now? Our advisors run this playbook with you, on your side of the table.
Google Cloud Advisory → Vendor Negotiation →
Run a software spend health check against your Google Cloud estate in under five minutes.
Open the Tool →
Deep Library

More on this topic.

Google Cloud Advisory →
When to Start a Google Cloud Negotiation and How to Sequence Pressure
Google Cloud · Guide
When to Start a Google Cloud Negotiation and How to Sequence Pressure
The full guide this article belongs to.
Guide
Does Google Cloud Discount More at Quarter End or Year End?
Google Cloud · Deep dive
Does Google Cloud Discount More at Quarter End or Year End?
Another angle on the same decision.
Guide
Google Cloud contract terms, and how to negotiate them
Google Cloud
Google Cloud contract terms, and how to negotiate them
Google Cloud discounts hinge on committed use and a custom agreement, not the list price.
Guide
Negotiating Oracle ERP Cloud pricing. A CIO playbook.
Google Cloud
Negotiating Oracle ERP Cloud pricing. A CIO playbook.
How to negotiate Oracle ERP Cloud pricing in 2026. The hosted named user metric, ramp deal
Guide
Vertex AI and Gemini, priced like infrastructure.
Google Cloud
Vertex AI and Gemini, priced like infrastructure.
Vertex AI and Gemini bill per token against a Google Cloud commit. Burn data, model routin
Guide
Editorial boardroom interior

The advisor your vendors do not want.

500+ enterprise clients. 11 vendor practices. Industry recognized. One conversation can change what you pay for the next three years.

Stay ahead of Google licensing changes.

One buyer side briefing a week. Renewal signals, audit moves, and the levers that work. No vendor spin.