Production token spend ran 2x to 5x above the pilot forecast at twelve months, and retrieval traffic rather than generation was the main driver
The model comparison is the easy part. The commitment you sign is priced against a forecast built from a pilot that never saw production traffic.
Prepared by Redress Compliance · August 18, 2026 · Enterprise GenAI vendor engagements. 18 to 26 engagements benchmarked, 2024 to 2025.
Executive summary
Pilot token estimates missed production volume by 2x to 5x, with retrieval and reranking traffic the main driver rather than the generation everybody models.
Private cloud and dedicated capacity carried a 30 to 60 percent premium over the shared interface for the same workload. Sovereignty is a real differentiator and it has a real price.
Routing search to the retrieval models rather than to full generation cut unit cost 20 to 40 percent on retrieval heavy use. The model mix is a cost lever before the rate card is.
Firms that committed early to the shared interface later paid a 30 to 60 percent premium to retrofit private deployment for data residency they had not planned for.
What is actually on offer here?
A narrower product line than the frontier vendors, by design. Three model families covering generation, embedding and reranking, aimed at enterprise retrieval and grounded generation rather than general purpose reasoning.
| Family | Purpose | Typical use | Pricing unit |
|---|---|---|---|
| Command R | Generation, retrieval augmented, tool use | Enterprise chat and summarization | Per million tokens |
| Command R+ | Larger generation, agentic | Complex retrieval and multi step tasks | Per million tokens at a higher rate |
| Embed v3 | Vector embeddings | Search, retrieval, clustering | Per million tokens at a low rate |
| Rerank v3 | Search result reranking | Retrieval quality and relevance | Per search query |
Retrieval first, not frontier first
The embedding and reranking products are the strongest part of the line, and the generation model is built specifically for retrieval pipelines. That shapes both the fit and the bill, and the vendor documents it on its product site.
Why does the deployment shape matter more than the rate?
Because sovereignty, data isolation and operational responsibility all sit downstream of it. Four shapes exist: the vendor managed interface, two hyperscaler catalogues, and private deployment inside your own environment.
Private deployment is the largest single differentiator against the frontier vendors. The models run inside the customer cloud or data centre, the data never leaves the perimeter, and the weights stay under vendor licensing.
It also carries a higher commitment and a managed service component. That is the 30 to 60 percent premium, and it is far cheaper to plan for than to retrofit after committing to the shared interface.
The GenAI consumption cost control brief
How to size a token commitment against measured drawdown rather than an adoption curve.
Get the brief →What 18 to 26 GenAI engagements showed
Across roughly 18 to 26 enterprise GenAI vendor engagements benchmarked between 2024 and 2025, including this one, production token spend ran 2x to 5x above the pilot forecast at twelve months. Three patterns recur.
- Forecasts understate consumption: pilot token estimates missed production volume by 2x to 5x, with retrieval and reranking traffic the main driver.
- Deployment choice moves price: private cloud and dedicated capacity carried a 30 to 60 percent premium over the shared interface for the same workload.
- Model family mix matters: routing search to reranking and embedding rather than full generation cut unit cost 20 to 40 percent on retrieval heavy use.
The pilot measures the question you asked it. Production measures every retrieval call underneath the answer, which is where the multiple comes from.
- Your agreements decoded into plain English before the auditor interprets them for you
- Coverage grid: liability caps, intellectual property protections and service levels checked in one pass
- A defensible position paper generated in minutes rather than weeks
How does the pricing actually work?
Token based for the generation and embedding models, per query for reranking, with a volume commitment discount on the enterprise tier. The published anchors are the floor rather than the price.
- Generation lists at $0.50 per million input tokens and $1.50 per million output.
- The larger generation model lists at $2.50 input and $10.00 output per million.
- Embedding lists at $0.10 per million tokens, and reranking at $2.00 per thousand searches.
Commitment bands, and what they buy
A $100K annual commitment carried 5 to 10 percent off list, $500K carried 10 to 20 percent, and $1M and above carried 20 to 30 percent. Those bands are worth less than getting the forecast right, because the discount applies to a number you chose.
Watch the briefing · 3:52Estimating the CommitmentHow to size an AI token commitment against measurement rather than an adoption curve.
Where does this vendor genuinely fit?
Where sovereignty and data isolation carry real weight, and where the workload is retrieval and classification rather than frontier reasoning. That is a narrower fit than the marketing suggests and a stronger one where it lands.
The contract carries the sovereignty, not the product
Deployment inside a hyperscaler catalogue keeps the models under vendor licensing while the tenancy is yours, documented for the Bedrock catalogue among others.
Enterprise data terms are the other half of the case. Data isolation is the default and there is no training on customer data in the enterprise terms, which is a contract position rather than a product feature.
The comparison against the frontier vendors, and the token economics underneath all of them, sit in the consumption billing guide, the pricing report and the frontier vendor comparison.
The clause level work is the same across every AI agreement, and it is set out in the contract red lines guide and the Claude enterprise guide.
What the engagements measured, 2024 to 2025
Two cuts of the engagement file describe the forecast risk and the structural one.
Measured at twelve months, with retrieval and reranking traffic rather than generation driving the gap.
Over the shared interface for the same workload, and the same premium paid again by anybody retrofitting it later.
The first number is a modelling problem you can fix before signing. The second is a sequencing problem, and it only has one cheap moment.
Your first five moves
- Model production traffic rather than pilot traffic, counting every retrieval and reranking call underneath the answer, since that is where the 2x to 5x gap forms.
- Decide the deployment shape before the commitment, not after, because retrofitting private deployment for data residency cost the same 30 to 60 percent premium a second time.
- Route search to the retrieval models rather than to full generation, worth 20 to 40 percent of unit cost on retrieval heavy workloads.
- Size the commitment band against the corrected forecast, because a 20 to 30 percent discount on an overstated volume is worse than no discount on the right one.
- Fix the data and intellectual property terms in the agreement rather than relying on the default. The GenAI practice prices the commitment against measurement before it is signed.
Frequently asked questions
How far off are pilot forecasts?
By 2x to 5x at twelve months across the engagements benchmarked. Retrieval and reranking traffic drove the gap, because a pilot measures the questions asked and production measures every call underneath the answer.
What does private deployment cost?
A 30 to 60 percent premium over the shared interface for the same workload, plus a higher commitment and a managed service component. It buys data residency the shared interface cannot offer.
Can private deployment be added later?
Yes, and it cost the same 30 to 60 percent premium to retrofit. Firms that committed early to the shared interface paid it a second time when residency requirements arrived.
How much does the model mix save?
Between 20 and 40 percent of unit cost on retrieval heavy workloads, by routing search to the embedding and reranking models rather than running everything through full generation.
What are the published rates?
Generation lists at $0.50 per million input tokens and $1.50 output, the larger model at $2.50 and $10.00, embedding at $0.10 per million, and reranking at $2.00 per thousand searches.
What do the commitment bands buy?
A $100K annual commitment carried 5 to 10 percent off list, $500K carried 10 to 20 percent, and $1M and above carried 20 to 30 percent. The band matters less than the volume it applies to.
Where does this vendor fit best?
Where sovereignty and data isolation carry weight and the workload is retrieval and classification. The line is deliberately narrower than the frontier vendors and stronger inside that scope.
What do the enterprise data terms say?
Data isolation is the default and there is no training on customer data. That is a contract position, which means it belongs in the agreement rather than in a product description.
Is the smaller model line a problem?
Only if you need frontier reasoning. For grounded generation over your own documents the smaller footprint is an operational advantage, particularly in private deployment.
What is the single biggest risk?
The forecast. Every commitment, discount band and deployment decision is priced against a number the pilot produced, and that number was wrong by 2x to 5x in the engagements measured.