Contents
Key takeawaysHow Gemini is pricedWhat we have seenAI spend and the commitModel routing savingsProvisioned throughputWhat lowers Google's pricingContract terms to ask forStopping spend driftWhat to do nextFAQVertex AI and Gemini spend is token and throughput pricing layered onto a Google Cloud commit. The deal turns on which meters you commit to and which you leave on demand.
- Two meters dominate. On demand token pricing and provisioned throughput drive most Vertex AI and Gemini spend.
- AI spend fills the commit. Vertex consumption draws down Google Cloud commitments, so AI growth can earn a better platform discount without padding the commit.
- Throughput cuts both ways. Provisioned throughput stabilizes cost and latency for production loads, and it becomes shelfware when bought ahead of demand.
- Routing beats rate cuts. Sending routine traffic to smaller Gemini models usually saves more than any rate discount Google will sign, and it needs no contract change.
- Bring the other two clouds. Documented OpenAI and AWS Bedrock quotes move Google's AI pricing because the three providers price against each other.
- Burn data beats forecasts. Committing to forecast AI growth repeats the classic cloud commit mistake with higher stakes, so commit only to what you have measured.
How are Vertex AI and Gemini priced in 2026?
Gemini on Vertex AI bills per token for on demand inference, with input and output priced separately for each model. Provisioned throughput, training and platform tooling run on their own meters, all published on the Vertex AI pricing page. On demand tokens and provisioned throughput drive most of the spend we review.
Google now markets the Vertex AI platform as Gemini Enterprise Agent Platform. The rename changes nothing in the commercial mechanics, and invoices, quotes and many documentation pages still use the Vertex AI name, so we use it here too.
| Model | Input | Output | Output to input ratio |
|---|---|---|---|
| Gemini 3.1 Pro Preview | $2.00 | $12.00 | 6 times |
| Gemini 2.5 Pro | $1.25 | $10.00 | 8 times |
| Gemini 2.5 Flash | $0.30 | $2.50 | About 8 times |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 4 times |
Output tokens cost 4 to 8 times as much as input tokens at current list prices, so workload shape matters as much as volume. A job that reads long documents and writes short answers costs far less per request than a drafting tool that writes pages from a short prompt.
Which pricing details change the bill most?
- Long context. On the Pro models, a prompt above 200,000 input tokens puts the whole request on a higher rate. Gemini 2.5 Pro goes to $2.50 input and $15 output per million tokens.
- Batch. Work submitted through the Batch API bills at 50 percent of the standard token price. Overnight extraction and classification jobs rarely need a real time answer.
- Priority tier. Google sells a priority tier for latency sensitive traffic at roughly 1.8 times the standard rate, for example $3.60 input and $21.60 output on Gemini 3.1 Pro Preview.
- Platform meters. Training, pipelines, vector search and the rest of the platform tooling each bill independently, and they rarely appear in the first AI proposal.
Treat published rates as the ceiling. What you actually pay depends on your commit, the discount instruments you choose and the alternatives you can put on the table.
Negotiating Google 3: The Cloud Bill and the AI Bill
What have we seen in recent Vertex AI and Gemini negotiations?
AI line items grew faster than any other Google Cloud meter and received the least scrutiny. That is the common thread across roughly 12 to 18 Google Cloud negotiations with material Vertex AI and Gemini components that I advised in 2024 and 2025.
- Forecasts ran hot. Token spend forecasts in first proposals overshot measured first year burn by 50 to 100 percent in most deals we reviewed.
- Routing cut unit cost. Moving routine workloads to smaller Gemini models cut effective token cost by 40 to 70 percent, with no change to the contract.
- Reserved capacity idled. Provisioned throughput bought ahead of measured load sat materially idle for the first two quarters in roughly half of the environments we reviewed.
Treat those ranges as benchmarks for what disciplined buyers achieved against the same Google sales approach. They are not promises. Your own meter history sets the baseline, and the first job in any Google AI negotiation is to produce it.
Vertex AI and Gemini negotiation guide
How to price tokens, savings plans, model tiers and fine tuning spend before you sign.
Get the white paper →How does Vertex AI spend fit a Google Cloud commit?
Vertex AI consumption draws down your Google Cloud spend commitment, so the AI negotiation is a commit negotiation. That works in your favor. Growing AI usage can fill a platform commit you already hold and justify a better discount across all Google Cloud services.
Two instruments matter. The first is the enterprise commit in your Google Cloud agreement, where private pricing on specific SKUs is negotiated. The second is the set of committed use discount mechanics that apply to the rest of the platform, which Google has extended to generative AI.
What do Flexible Savings Plans do for Gemini spend?
Flexible Savings Plans are Google's spend based commitments for generative AI. They cover first party Gemini models, open source models sold as a service and participating third party models in Model Garden. A 1 year plan gives a 10 percent discount and a 3 year plan gives 20 percent, measured in dollars per month.
Read the terms before you sign one. You cannot cancel a plan, usage above the monthly amount bills at the on demand price, and provisioned throughput subscriptions and some fine tuning services draw down the plan without earning any extra discount. A 3 year plan also fixes your spend level while model prices keep falling.
How should you structure the AI component?
- Measure current token and throughput burn by model and workload before any commitment conversation.
- Commit platform wide at or below cleaned trailing spend and let AI growth fill the commit. See our guide to Google Cloud committed use discounts for how the tiers work.
- Keep flexibility at the model level. Never commit to a single model family in a market that reprices every quarter.
- Review provisioned throughput every quarter against measured load, never against launch forecasts.
How much can model routing cut Vertex AI token spend?
Routing is usually the largest single saving, and it needs no contract change. The price gap between Gemini tiers is wider than any discount Google will sign, so sending each workload to the cheapest model that passes your quality tests beats most rate negotiations.
Say a company sends 20 billion input and 4 billion output tokens a month, all to Gemini 2.5 Pro. At list price that is $25,000 for input and $40,000 for output, or $65,000 a month. Now route 40 percent of the traffic to Flash-Lite, 30 percent to Flash and keep 30 percent on Pro.
| Model | Share of traffic | Input cost | Output cost | Monthly total |
|---|---|---|---|---|
| All traffic on 2.5 Pro | 100 percent | $25,000 | $40,000 | $65,000 |
| 2.5 Pro after routing | 30 percent | $7,500 | $12,000 | $19,500 |
| 2.5 Flash after routing | 30 percent | $1,800 | $3,000 | $4,800 |
| 2.5 Flash-Lite after routing | 40 percent | $800 | $640 | $1,440 |
| Routed total | 100 percent | $10,100 | $15,640 | $25,740 |
The routed bill is $25,740 a month, about 60 percent lower, and the annual difference is roughly $471,000. Moving the overnight Flash-Lite jobs to the Batch API would cut that line by another half.
When is provisioned throughput worth buying for Gemini?
Provisioned throughput is worth buying only for measured production load that needs stable latency and a predictable cost. Bought ahead of demand, it becomes shelfware, because you pay for the reserved capacity whether traffic arrives or not.
Google sells it in generative AI scale units (GSUs), each a fixed slice of model throughput, on four term lengths. From July 1, 2026, regional endpoints cost about 10 percent more than the global rates below for generally available Gemini 3 and later models.
| Term | Price per GSU | Commitment for 20 GSUs |
|---|---|---|
| 1 week | About $1,200 per week | About $24,000 for the week |
| 1 month | $2,700 per month | $54,000 |
| 3 months | $2,400 per month | $144,000 |
| 1 year | $2,000 per month | $480,000 |
The annual term looks cheapest, but idle capacity erases the gap. Twenty GSUs on a 1 year term cost $40,000 a month. If half sit idle for two quarters, $120,000 buys nothing. Ten GSUs on 1 month terms would have cost $27,000 a month over the same period, $13,000 less, with room to add units as load grows.
Why we disagree with locking in throughput early
The usual advice is to reserve provisioned throughput early because AI capacity is scarce and prices only rise. We disagree. In the 12 to 18 Google Cloud AI negotiations I advised in 2024 and 2025, per token prices for equivalent capability fell repeatedly as new Gemini tiers shipped. Early reservations sat half idle while better models arrived at lower rates.
Reserve only for measured production load, keep the right to shift reserved capacity to newer models, and let falling prices work for you. Scarcity arguments are a sales tactic, so ask Google to show the constraint in writing for your region. Our note on capacity scarcity in Google Cloud deals covers how to test the claim.
Commit to the burn you can measure. The market keeps lowering the price of every model you have not bought yet.
What lowers Google's AI pricing in a negotiation?
Three things lower Vertex AI and Gemini pricing: documented competitor quotes, measured burn data in place of growth forecasts, and routing discipline that proves you control consumption. Google's Cloud terms leave enterprise AI pricing open to negotiation inside your agreement. A quarterly throughput review clause adds a fourth tactic on the capacity side.
| Tactic | Works when | Typical result |
|---|---|---|
| OpenAI and AWS Bedrock quotes on the table | Current, written and matched to the workload | Resets the AI rate conversation |
| Burn data in place of forecasts | Twelve months of meter history | Cuts committed AI volume 30 to 50 percent |
| Model routing discipline | Routine traffic runs on smaller models | 40 to 70 percent off token spend |
| Throughput right sizing | A quarterly review clause in the contract | Removes idle reserved capacity |
Why rival cloud quotes move Gemini rates
Google, Microsoft and AWS price frontier AI against each other week by week. A written quote from either rival, matched to your workload, shifts a Gemini rate faster than any volume argument. Get it for the same task, token mix and latency, using our guides to AWS Bedrock pricing and Azure OpenAI negotiation.
What will the Google account team say, and how should you answer?
- "Add your AI growth to the commit and we can move you up a discount tier." Answer that you will commit at trailing spend and let growth fill it, and ask what tier your measured spend earns today.
- "Capacity is tight, so reserve throughput now." Ask for the constraint in writing for your region and model, then reserve only your measured production peak on a short term.
- "A 3 year savings plan gets you 20 percent." Point out that the plan cannot be canceled while model prices fall every few months, and ask what the 1 year rate becomes with a model substitution right.
- "Our list price already beats OpenAI." Put the matched quote on the table and ask Google to price the same workload at the same token mix.
Which contract terms should a Vertex AI agreement include?
Ask for terms that keep your commitment flexible as models change. Each one below addresses a cost problem described earlier on this page. Put them in the enterprise agreement itself so they survive changes in the Google account team.
- Drawdown scope. A written list confirming that Gemini tokens, partner models in Model Garden, provisioned throughput and fine tuning all count toward the commit, so no AI line falls outside it.
- Model substitution. The right to apply commitments and reserved capacity to newer Gemini models at no penalty. Our swap and reallocation clause guide has sample wording.
- Price protection on successors. A new model tier that replaces one you use should cost no more per token than the tier it replaces, for the length of the term.
- Quarterly throughput review. The right to reduce or convert GSUs each quarter against measured load.
- Shortfall handling. Unused commit rolls into the next period, never billed as a penalty. The commit clause redlines guide covers the language.
How do you stop AI spend drifting after signature?
Set routing policy and quota governance at signature, before the first surprise invoice. AI spend drifts because every team can call a frontier model by default, and no contract term controls that.
- Default to small. Route routine classification and extraction to the cheapest Gemini tier that passes your quality checks.
- Quota by team. Give each project a budget and alerts on token meters from day one.
- Quarterly model review. Reprice workloads as new model tiers ship, because last year's routing table overpays today.
How do you check your own Vertex AI consumption?
Export Cloud Billing data to BigQuery and filter it by service and SKU. That gives you spend per model, per project and per meter, the twelve month history Google will ask for. Cloud Monitoring shows token counts and throughput usage per model, which you need to size any GSU purchase.
Cloud Billing budgets alert you but do not cap spending on their own. To stop a runaway project you need a programmatic response to the budget notification, or per project quotas. Our guide to Google Cloud cost allocation and tagging explains how to label AI workloads so each team sees its own bill.
What to do next
- This month. Export twelve months of Vertex AI and Gemini meter data by model and project.
- Before the first commit meeting. Build a routing table that maps each workload to the cheapest model tier that passes your quality tests.
- Before any new order. Right size or cancel provisioned throughput that idles below measured production load.
- Three months before renewal. Collect current written quotes from OpenAI and AWS Bedrock for matched workloads.
- At the commit. Fold cleaned AI burn into the platform commit at or below trailing consumption, with drawdown scope and model substitution written in.
- After signature. Set project level token quotas and a quarterly model repricing review. Our guide to negotiating Google Cloud AI contracts covers the wider agreement.
Frequently asked questions
How is Gemini priced on Vertex AI?
Per million input and output tokens at on demand rates on the Vertex AI pricing page, with a separate rate for each model and a higher rate for Pro prompts above 200,000 tokens. Enterprise agreements then apply negotiated rates or savings plans against committed Google Cloud spend.
Does Vertex AI spend count toward a Google Cloud commit?
Yes. Vertex consumption draws down platform commitments, which makes it the main point of negotiation. Confirm in writing that partner models, provisioned throughput and fine tuning are also eligible, because Google treats some of these SKUs differently in its savings plans.
Is provisioned throughput worth buying for Gemini?
Only for measured production load that needs stable latency and a predictable cost. Size the GSU count from Cloud Monitoring peaks, start on 1 month terms, and move to a 1 year term only after two or three months of steady traffic on the same model.
What is the fastest way to cut Vertex AI spend?
Model routing. Move routine classification, extraction and summarization traffic to smaller Gemini tiers, then push work that can wait overnight to the Batch API at half the standard token price. Neither step needs Google's agreement or a contract amendment.
Do OpenAI quotes really move Google's pricing?
Yes, when the quote is current, written and matched to your workload. Google, Microsoft and AWS price enterprise AI against each other, so a like for like Azure OpenAI or Bedrock price gives the account team a reason to escalate for a better Gemini rate. A vague mention of a competitor does little.
Should we commit to forecast AI growth?
No. Forecasts in first proposals overshot measured first year burn by 50 to 100 percent in our engagement file. Size the commit on trailing consumption, since a commit is far easier to raise mid term than to reduce.
Is a Google Cloud Flexible Savings Plan worth it for Gemini?
Sometimes, for a stable base of token spend. The 1 year plan saves 10 percent and the 3 year plan 20 percent, but neither can be cancelled. Size it below your lowest measured month so falling model prices do not leave you paying for spend you no longer need.