Contents
Key takeawaysWhat Cohere sellsHow API pricing worksWhy production outruns the pilotChoosing a deployment optionWhat we have seenWhere Cohere fitsWhat the account team will sayContract terms to ask forChecking your real usageWhat to do nextFAQCohere's list rates are the easy part. The commitment you sign is priced against a pilot forecast that usually misses production spend by a multiple, and the deployment shape you pick sets much of the rest.
- Count what sits under each answer. Production token spend ran 2x to 5x above pilot forecasts at twelve months, driven by retrieval and reranking calls.
- Choose deployment before commitment. Private and dedicated capacity cost more than the shared API for the same workload, and buyers who retrofitted it later paid that premium twice.
- Route lookups to Embed and Rerank. Sending search traffic to the retrieval models cut unit cost 20 to 40 percent on retrieval heavy work.
- Discount bands follow the forecast. The deeper commitment bands only pay off when the production volume behind them is real.
- Check the model version. Older Command R and Command R+ versions were retired on September 15, 2025, so ask for successor pricing in the contract.
- Sign the data terms. Put no training, the retention period and zero data retention in the order form instead of relying on dashboard settings.
Cohere usually reaches a GenAI shortlist on retrieval quality, private deployment or data residency. Comparing its models against OpenAI or Anthropic is the easy part. The harder part is the commitment, because it is priced against a token forecast from a pilot that never carried production traffic.
What does Cohere sell to enterprises?
Cohere sells a narrower product line than the frontier vendors, by design. Three model families cover generation, embedding and reranking, and they are aimed at enterprise retrieval and grounded generation over your own documents rather than general purpose reasoning.
| Family | Purpose | Typical use | Pricing unit | Current versions |
|---|---|---|---|---|
| Command R | Generation, retrieval augmented, tool use | Enterprise chat and summarization | Per million tokens | Command R (08-2024); Command A for new work |
| Command R+ | Larger generation, agentic | Complex retrieval and multi step tasks | Per million tokens at a higher rate | Command R+ (08-2024); Command A family |
| Embed v3 | Vector embeddings | Search, retrieval, clustering | Per million tokens at a low rate | Embed 4, with v3 still available |
| Rerank v3 | Search result reranking | Retrieval quality and relevance | Per search query | Rerank 3.5, Rerank 4 Fast and Rerank 4 Pro |
Command A is now the flagship generation model, with Reasoning, Vision and Translate variants. Cohere also sells North, a workplace assistant, and Compass, an enterprise search product. Both are quoted by sales at custom enterprise pricing, with no public rate card.
Why the retrieval models carry the case
The embedding and reranking products are the strongest part of the line, and the Command models are built for retrieval pipelines. That shapes both the fit and the bill. Cohere documents the line on its product site, and the model pages list which clouds carry each version.
How does Cohere API pricing work?
Generation and embedding are billed per token, reranking per search, and the enterprise tier adds a volume commitment discount. The published rates are the starting point for a negotiation, and the price you pay is set by the commitment band and the terms around it.
| Model | Input | Output | Status |
|---|---|---|---|
| Command R (03-2024) | $0.50 per million tokens | $1.50 per million tokens | Deprecated September 15, 2025 |
| Command R (08-2024) | $0.15 per million tokens | $0.60 per million tokens | Live |
| Command R+ (08-2024) | $2.50 per million tokens | $10.00 per million tokens | Live |
| Embed v3 | $0.10 per million tokens embedded | Available; Embed 4 now current | |
| Rerank v3 | $2.00 per thousand searches | Rerank 3.5 and Rerank 4 now current | |
Two details change the forecast. Cohere counts one rerank search as one query with up to 100 documents, so a pipeline that reranks several hundred candidates per question pays for several searches. And the pricing page now leads with North, Compass and Model Vault, so current Command A rates come from your account team or the cloud marketplace listing.
What do the commitment bands buy?
In the deals we benchmarked, a $100K annual commitment carried 5 to 10 percent off list. A $500K commitment carried 10 to 20 percent, and $1M and above carried 20 to 30 percent. Those bands are worth less than getting the forecast right, because the discount applies to a volume you chose.
Trial keys, production keys and rate limits
Check which key your pilot ran on, because Cohere's documented limits differ sharply by key type:
- Trial key. Free, capped at 1,000 API calls a month and 20 chat requests per minute.
- Production key. Pay as you go, with 500 chat requests per minute on established models, 1,000 rerank requests per minute and 2,000 embed inputs per minute.
- Newer models. Production limits for the Command A variants are set by Cohere sales, so ask for them in writing.
Why does Cohere production spend run so far above the pilot forecast?
Because the pilot forecast counts the question and its answer, and production pays for everything underneath. Each question triggers a query embedding, a rerank over the candidate documents and a generation call whose input is mostly retrieved text. Retrieval and reranking traffic drove most of the gap we measured at twelve months, more than the generation step most teams model.
A worked example
Say you run an internal knowledge assistant answering 20,000 questions a working day, or 440,000 a month, on Command R+ (08-2024). The pilot team priced 500 input tokens and 400 output tokens per question. In production each prompt also carries six retrieved passages, taking input to 4,000 tokens, and every question is reranked once.
| Cost line per question | Pilot forecast | Production | Production with routing |
|---|---|---|---|
| Input tokens | 500 = $0.00125 | 4,000 = $0.01000 | $0.01000 on 60 percent of questions |
| Output tokens | 400 = $0.00400 | 400 = $0.00400 | $0.00400 on 60 percent of questions |
| Rerank | Not counted | 1 search = $0.00200 | 1 search = $0.00200 on every question |
| Monthly total | $2,310 | $7,040 | $4,576 |
| Annual total | $27,720 | $84,480 | $54,912 |
Production comes out at about 3 times the pilot number, inside the range we see, before any growth in users. Embedding the queries, at about 50 tokens each, adds roughly $2 a month here. Embedding the whole document library again after a model change costs far more, so budget it as a separate line.
The routing column assumes 40 percent of questions are lookups, answered by embedding and rerank with no generation call. Those questions drop from $0.016 to $0.002 each, cutting the monthly bill by $2,464, or 35 percent. That is inside the range we measured on retrieval heavy work, set out in the evidence section below.
Why a larger band can cost more
At $84,480 a year, the account team will suggest the $100K band. Say routing brings real consumption to $54,912 at list. With 10 percent off, you draw down about $49,400, so a take or pay commitment leaves roughly $50,600 unused. Paying as you go would have cost $54,912. Size the band on the corrected, routed forecast.
Which Cohere deployment option should you choose?
Choose the deployment shape before you choose the commitment, because data isolation, sovereignty and who runs the infrastructure all follow from it. Cohere sells through four routes, and each one changes who bills you and where your prompts go.
- Cohere SaaS API. Shared infrastructure, billed per token by Cohere. Logged prompts and generations are deleted after 30 days, and zero data retention is available to enterprise customers who make additional usage commitments.
- Cloud marketplaces. Cohere models run in Amazon Bedrock and SageMaker, Microsoft Azure AI Foundry and Oracle OCI Generative AI. The tenancy is yours, the models stay under Cohere licensing, and Cohere says it receives no prompts or outputs. The Bedrock catalogue is one example, and not every model version is on every cloud.
- Model Vault. Dedicated capacity run by Cohere, billed by the hour or on monthly and annual terms. Published examples include Embed 4 Small at $4.00 an hour or $2,500 a month, and Rerank 4 Pro Large at $10.00 an hour or $6,500 a month.
- Private deployment. The models run inside your own cloud account or data center, the data never leaves your perimeter, and the weights stay under Cohere licensing.
Private deployment is Cohere's clearest difference from the frontier vendors. It also brings a higher commitment and a managed service component. In our benchmarks, private and dedicated capacity carried a 30 to 60 percent premium over the shared interface for the same workload.
Why retrofitting private deployment costs you twice
Firms that committed early to the shared interface paid that premium again when residency rules forced a rebuild on private deployment. Unless the contract says otherwise, spend committed on the shared API does not follow you to the new deployment. If residency is plausible within the term, price private deployment now, or write a migration right into the first agreement.
What have we seen in recent Cohere and GenAI negotiations?
We benchmarked roughly 18 to 26 enterprise GenAI vendor engagements between 2024 and 2025, including Cohere deals. Measured at twelve months, production token spend came in at 2x to 5x what the pilot had predicted. Three patterns came up again and again:
- Forecasts understate consumption. Pilot token estimates missed production volume, and retrieval and reranking traffic was the main driver.
- Deployment choice changes the price. Private cloud and dedicated capacity priced well above shared API access for identical work, and retrofitting it cost the premium a second time.
- The model mix matters. Routing search to reranking and embedding instead of full generation cut unit cost 20 to 40 percent on retrieval heavy use, and it did so before any rate card discount.
The pilot measures the question you asked it. Production measures every retrieval call underneath the answer, and that is where the multiple comes from.
Why choosing the model first is the wrong order
The usual advice is to run a model bake off, pick the winner on quality and then negotiate the rate. We think that order gets the cost wrong.
Rate differences between shortlisted models are small next to a forecast that is off by several times, or a deployment choice you have to reverse. Model production traffic and settle the deployment shape first, then compare rate cards against that volume.
Where does Cohere fit, and where does it not?
Cohere fits where sovereignty and data isolation carry real weight and the workload is retrieval, search and classification. It is a weaker choice for frontier reasoning. That is a narrower fit than the marketing suggests, and a stronger one where it lands.
Put the data terms in the agreement
Cohere's enterprise data commitments apply to commercial, paying customers and are built around data isolation. On the SaaS platform, though, training use of prompts and generations is an opt out setting under Data Controls in the dashboard. Get no training written into the agreement, where an admin cannot switch it back.
The comparison with the frontier vendors, and the token economics under all of them, are in our consumption billing guide, the pricing report and the frontier vendor comparison. The clause work is the same across AI agreements, set out in the contract red lines guide and the Claude enterprise guide.
What will the Cohere account team say, and how should you answer?
- "Your pilot usage supports the $100K band." Ask them to price the band against a production model that counts rerank searches and retrieved context, and agree the forecast before the discount.
- "Start on the API and move to private deployment later." Ask what share of the committed spend transfers if you move, and put the answer in the order form.
- "The enterprise terms already cover data." Ask for no training, the retention period and zero data retention, if you need it, written into your agreement.
- "Command R is the older generation, so move to Command A." Ask for Command A rates in writing, the successor model pricing and a migration window before the old version is retired.
What should a Cohere enterprise contract include?
Ask for terms that protect you against a forecast that turns out wrong in either direction, and against model changes during the term:
- Price hold on every model family. Command, Embed and Rerank rates fixed for the full term, so a change in mix does not reprice you.
- Commitment that spans families. Spend on embedding, reranking and generation draws down one commitment, so routing savings stay yours.
- Successor pricing. A deprecated model's rate carries to its replacement, since Command R (03-2024) and Command R+ (04-2024) were retired on September 15, 2025.
- Rollover or ramp. Unused commitment carries into the next year, or the commitment steps up with measured adoption.
- Deployment migration credit. Spend on the SaaS API counts toward a later private or Model Vault deployment.
- Data terms. No training, a stated retention period, zero data retention where needed and deletion on exit.
How do you check your real Cohere usage before you commit?
Measure from billing data. Every Cohere API response returns a billed units block with input tokens, output tokens and search units. Log it per call and per use case, and you have the real mix of generation, embedding and reranking.
- Cohere dashboard. The billing and usage pages show spend, and the API keys page shows whether the pilot ran on a trial or production key.
- Amazon Bedrock. CloudWatch reports invocations and input and output token counts per model, and Cost Explorer shows the spend.
- Azure and Oracle OCI. Cost management reports per deployment on each cloud's own invoice.
Run the production pipeline on a sample of real traffic for a few weeks, then scale by expected users. Size the commitment on that number. For the spend side, see our token usage cost surge report.
What to do next
- Model production traffic. Count every retrieval, rerank and embedding call under the answer, since that is where the gap between pilot and production forms.
- Decide the deployment shape first. Settle SaaS, marketplace, Model Vault or private deployment before the commitment, because a later retrofit pays the premium again.
- Route search to the retrieval models. Send lookups to embedding and rerank instead of full generation, and measure the saving on a sample.
- Size the band on the corrected forecast. A 20 to 30 percent discount on an overstated volume costs more than no discount on the right one.
- Write the data and intellectual property terms into the agreement. Do not rely on dashboard defaults or published policy.
- Get the commitment priced against measurement. Our GenAI practice reviews the forecast and the terms before you sign.
Frequently asked questions
How far off are Cohere pilot forecasts?
By 2x to 5x at twelve months in the engagements we benchmarked. Pilots often run on trial keys with a small document set and a few users, so they miss the reranking calls, the retrieved context in every prompt and the growth in usage after rollout.
What does Cohere private deployment cost?
In our benchmarks, a 30 to 60 percent premium over the shared interface for the same workload, plus a higher commitment and a managed service component. You also supply the GPU capacity in your own cloud account or data center, which is not in Cohere's quote. In return you get data residency the shared API cannot offer.
Can Cohere private deployment be added later?
Yes, but buyers who committed to the shared API first paid the private deployment premium again when residency rules arrived. If residency is a realistic prospect within the term, negotiate a migration credit now so that SaaS spend counts toward the later deployment.
How much does routing between Cohere models save?
Enough on retrieval heavy workloads to change which commitment band makes sense. A lookup answered by Embed and Rerank costs a fraction of a Command generation call, so the saving depends on what share of your traffic needs a written answer. Measure that share on real traffic before sizing the commitment.
What are Cohere's published API rates?
Command R (03-2024) listed at $0.50 per million input tokens and $1.50 output, and Command R+ (08-2024) at $2.50 and $10.00. Embed v3 lists at $0.10 per million tokens and Rerank at $2.00 per thousand searches. The current Command R version is cheaper, and Command A rates come from Cohere sales or the cloud marketplace.
What do Cohere commitment bands buy?
In the deals we benchmarked, discounts ran from 5 to 10 percent off list at a $100K annual commitment to 20 to 30 percent at $1M and above. The band matters less than the volume it applies to, so agree the production forecast first and ask whether unused commitment rolls over.
Where does Cohere fit best?
In retrieval, search, classification and grounded answers over your own documents, particularly where data must stay inside your cloud account or country. Workloads that need frontier reasoning may be better served by a frontier model, with Cohere kept for the retrieval layer underneath.
What do Cohere's enterprise data terms say?
They apply to commercial, paying customers and rest on data isolation. On the SaaS API you opt out of training under Settings, Data Controls, and zero data retention is granted on request only to customers who make extra usage commitments. In marketplace and private deployments Cohere says it receives no prompts or outputs.
Is Cohere's smaller model line a problem?
Only if you need frontier reasoning. For grounded generation over your own documents the smaller footprint is an operational advantage, because smaller models need less GPU capacity, and that matters most when you run them yourself in a private deployment.
What is the biggest risk in a Cohere contract?
The forecast. The commitment, the discount band and the deployment decision are all priced against a number the pilot produced, and in the engagements we measured that number was wrong by a multiple. Correct it with production traffic before any of the three is signed.