Sizing the commitment at p50 billed for 15 to 25 percent that never ran
A committed use discount trades rigidity for rate, so the deepest discount goes to the most specific commitment. That makes the sizing rule, not the rate, the decision that sets what you actually pay, and a forecast that is wrong half the time is a poor place to put the deepest commitment.
Prepared by Redress Compliance · August 15, 2026 · Google Cloud advisory. Based on 25 to 35 benchmarked Google Cloud estates, 2024 to 2025.
Executive summary
The most expensive mistake was not under discounting, it was over committing. Resource based commitments sized to the p50 forecast left 15 to 25 percent of the commitment unused but still billed.
The p20 rule fixes it: commit resource based only to the level you exceed 80 percent of the time, taken from twelve months of hourly telemetry rather than from a plan.
Layering beats depth. Resource based at the floor, flexible across the variable band, on demand for the burst, which moves 15 to 30 percent of compute spend in most estates.
Flexible commitments were routinely ignored in favor of deeper resource based rates, then stranded when workloads moved region.
BigQuery slots landed 30 to 40 percent below on demand when sized to steady state rather than to peak, which is the same discipline applied to a different meter.
The three variants, on one flexibility curve
| Variant | Discount depth | Flexibility | Where it belongs |
|---|---|---|---|
| Resource based | Deepest | Low: fixed machine family and region | The p20 floor only |
| Flexible | Moderate | High: applies across families and regions | The variable band above the floor |
| Spend based | Shallowest | Managed service spend | BigQuery and similar services |
| On demand | None | Total | The burst, deliberately |
The p20 baseline rule. Take twelve months of hourly consumption, classify it by machine family and region, and find the level you exceed 80 percent of the time. That is the p20 floor, and it is where the resource based commitment goes. Everything above it belongs to flexible commitments or on demand. The floor is almost never wrong, which is exactly why the deepest discount belongs there and nowhere else. Classification matters as much as the percentile: resource based commitments apply only to the family and region committed, so an unclassified commitment strands the moment a workload moves.
The levers that actually move the rate
- Architecture before rate: resource based for the floor, flexible for the variable band, on demand for the burst. The structure moves more money than the percentage does.
- Stagger the end dates across quarters so the estate never faces a single renewal cliff with the whole commitment portfolio as the stake.
- Stack the platform discount on top of the commitment architecture rather than treating the two as alternatives.
- Bring telemetry, not a forecast, because the commitment that survives the term is the one built from history, as the tactics brief sets out.
- Size BigQuery slots to steady state, not to the busiest hour, which is the same failure mode on a different meter.
- Keep a credible alternative live before signing, since it is the only lever that improves the rate without committing you to more volume.
The Google Cloud CUD negotiation framework
The CUD math, the p20 baseline rule, and the commitment sequence our advisors use on live engagements.
Get the framework →An unused deep discount costs more than a used shallow one
The account team pitch is to maximize the resource based commitment, because it carries the deepest rate. In roughly 25 of the 35 estates we benchmarked, that deeper rate was wiped out by 15 to 25 percent of unused commitment that billed anyway when the forecast missed. The instruction is not wrong about the rate. It is wrong about which number the rate applies to.
The reason p50 fails is definitional rather than bad luck. A p50 forecast is the level you exceed half the time, which means half the hours of the year fall below it by construction. Putting the most rigid, deepest, least recoverable commitment on a number designed to be wrong half the time guarantees paid idle capacity. p20 inverts the logic: commit where you are almost always above the line, and let the uncertain band ride on instruments that can absorb being wrong. The deep discount then applies to hours that genuinely exist.
Classification is the part teams skip, and it produces a quieter version of the same loss. Resource based commitments are tied to a machine family and a region, so a commitment that is correctly sized in aggregate can still strand when a workload moves to a different region or shifts shape. The estate then pays twice: the committed capacity in the old location and on demand rates in the new one. Twelve months of hourly consumption classified by family, region and service is what prevents that, and it is also the evidence that makes the negotiation itself defensible.
What follows is a different kind of negotiation. Instead of arguing about a percentage, you arrive with a commitment shape that is demonstrably safe, which changes what the discussion is about. The platform wide discount stacks on top and is where the remaining negotiation genuinely lives, and staggered end dates keep a comparison point in the market each quarter rather than concentrating all leverage on one date. The wider commitment picture sits in the CUD guide, coverage discipline in the enterprise playbook, and the ongoing rhythm in the FinOps playbook.
Watch the briefing · 6:33Google Cloud: Is There Leverage? Five TacticsA credible alternative is the only lever that improves the committed use rate without committing you to more volume, and it exists before signing and evaporates after.
- CUD, EDP and MACC commits sized from actual consumption, classified by family and region
- Right sizing for databases, storage and compute schedules with dollar figures
- A ranked savings queue your team can work through
What the benchmark file shows
Across roughly 25 to 35 Google Cloud estates benchmarked in 2024 and 2025, the sizing rule separated the good outcomes from the expensive ones:
On resource based commitments sized to the median forecast, which is wrong half the time by construction.
Resource based at the p20 floor, flexible across the variable band, on demand for the burst.
The patterns: commitments sized from forecasts rather than telemetry, flexible commitments ignored for a deeper rate, and consumption never classified by family and region until it stranded.
The buyer side move is to commit to the floor you can prove. The wider library sits in the Google Cloud practice.
Your first five moves
- Export twelve months of hourly consumption and classify it by machine family, region and service before anything is priced.
- Calculate the p20 floor and commit resource based only to that level.
- Cover the variable band with flexible commitments, and leave the burst deliberately on demand.
- Size BigQuery slot commitments to steady state, not to peak.
- Stagger end dates across quarters and stack the platform discount on top. The Google negotiation service builds the architecture with you.
Frequently asked questions
What is the p20 baseline rule?
Take twelve months of hourly consumption, find the level you exceed 80 percent of the time, and commit resource based only to that floor. Everything above it goes to flexible commitments or on demand. The floor is almost never wrong, so the deep discount is almost never wasted.
Why does committing at p50 cost money?
Because a p50 forecast is wrong half the time by construction. In the estates we benchmarked, resource based commitments sized to p50 left 15 to 25 percent of the commitment unused but still billed, which erased the advantage of the deeper rate the sizing was chosen to capture.
What are the three CUD variants?
Resource based commits to specific vCPU and memory in a region for the deepest but most rigid discount. Flexible commits to an hourly spend that applies across machine families and regions. Spend based commits to a managed service spend and is shallower again. They sit on a flexibility curve, and the deeper the rigidity the deeper the rate.
Why classify consumption by machine family and region?
Because resource based commitments only apply to the exact family and region committed. Classify twelve months of consumption first, or the discount strands the moment a workload moves region or changes shape, leaving you paying for capacity in one place and on demand rates in another.
How much does a layered architecture actually move?
In most estates, 15 to 30 percent of total compute spend. Resource based at the p20 floor, flexible across the variable band, and on demand for the burst. The saving comes from coverage discipline rather than from negotiating a deeper headline rate.
What about BigQuery slot commitments?
Sized to steady state rather than to peak, slot commitments landed 30 to 40 percent below on demand in the estates we reviewed. The failure mode is identical to compute: sizing to the busiest hour and paying for that capacity in every other hour of the year.
Should commitment end dates be staggered?
Yes. Spreading end dates across quarters avoids a single renewal cliff where the whole estate is exposed at once, and it means no negotiation ever happens with the entire commitment portfolio as the stake. It also keeps a live comparison point in the market each quarter.
Do CUDs replace the platform discount?
No, they stack. The platform wide discount layers on top of the commitment architecture rather than replacing it, so treating them as alternatives leaves value on the table. Model the commitment structure first, then negotiate the platform discount on top of it.
Negotiating Google 3: The Cloud Bill and the AI Bill
CUDs stack on private rates, the commit contract's three deciding clauses, support billed on list price, and the three layer AI bill with its double pay risk.