HomeTraining AcademyMicrosoft Agreements and CopilotSession 38
Microsoft Agreements and Copilot · Module 8 ยท Advanced situations and the capstone · Session 38 of 40 · 18:30

Copilot at enterprise scale

The rollout, the ramp, the consumption tail, and the renewal that follows the first full year. Three knowledge checks along the way, and 1 clip from a senior cloud advisor.

The presenter in this session is an AI generated avatar. The curriculum and guidance are real, produced by Redress Compliance analysts from our consulting engagements and market network.

What you will be able to do after this session

  • 1Scale changes the failure mode. At pilot size the risk is a wasted pilot. At estate size it is a locked commitment against an adoption curve that does not follow the plan.
  • 2The ramp is the design. Named tranches released on measured thresholds, with the rate held for full volume. That structure was worth 15 to 30 percent against the all in bundle.
  • 3The tail is the second bill. Agents, credits, and message meters grow after the seats stop growing, and they have no headcount ceiling at all.
  • 4The realised multiple. Fully loaded cost per active user ran 1.8 to 3.4 times the headline, and at scale that gap becomes the number the board sees.
  • 5The renewal is decided by year one. Estates arriving with documented weekly active usage negotiated from their own telemetry. Estates without it negotiated against an adoption narrative.

How the session works

This is a taught session, not a talking head. The instructor works through analyst grade slides, and three times the video stops on a question with four options on screen. Pause, commit to an answer, and the next slide explains which option is right and why each of the others is wrong. Once in the session the frame splits and a senior cloud advisor gives the view from inside real Oracle negotiations, and the instructor picks the clip apart when the slides return.

Homework before the next session, about an hour

  • 1Split adoption by persona. Not the estate average. Twelve curves inside one number, and the spread is the finding.
  • 2Plot seats against the metered line. Twelve months of each on one chart. If the second is rising while the first is flat, you have found the tail.
  • 3Check the exits. Was a rule agreed for tranches that miss their threshold, and has it ever been applied? Both answers are informative.
  • 4Compute cost per active user. Full stack over sustained weekly actives. Compare against 1.8 to 3.4 times headline and see where you sit.
  • 5List the roles that did not benefit. Write them down now, before the renewal, because that list is what makes the rest of your paper believable.

Session transcript

The full narration of this session, section by section, for reading and reference. Guest analyst clips are marked.

Welcome and objectives 0:02

Welcome back, session thirty eight of forty. Module four covered Copilot from the subscription to the business case. Today is what changes when a deployment goes properly large, and I want to open with the distinction that matters. At pilot size, the risk of getting Copilot wrong is a wasted pilot, which is embarrassing and cheap. At estate size, the risk is a material multi year commitment sitting against an adoption curve that did not follow the plan, in a line item large enough that the board is watching it. Those are different problems and they need different disciplines. And there is a second thing that only appears at scale, which is the consumption tail. The seats plateau, and then the metered layer keeps growing, on its own curve, about a year behind, in the gap between two owners. That second bill is the half most organisations do not see coming.

Five takeaways. One, scale changes the failure mode: at pilot size the risk is a wasted pilot, at estate size it is a locked commitment against an adoption curve that does not follow the plan. Two, the ramp is the design: named tranches released on measured thresholds with the rate held for full volume, which was worth fifteen to thirty percent against the all in bundle. Three, the tail is the second bill: agents, credits, and message meters grow after the seats stop growing, with no headcount ceiling at all. Four, the realised multiple: fully loaded cost per active user ran one point eight to three point four times the headline, and at scale that gap becomes the number the board sees. Five, the renewal is decided by year one, because estates with documented weekly active usage negotiated from their own telemetry and everybody else negotiated against a narrative.

What changes at scale 2:12

What changes at scale, three things that behave differently above a few thousand seats. Adoption stops being uniform: at pilot size you are measuring one population, and at estate size you have a dozen adoption curves hidden inside one average, which conceals both the roles that are thriving and the ones that never started. The consumption tail appears: seats plateau while the metered layer keeps growing, because agents get built after the rollout rather than during it, so the second bill arrives about a year after the first decision. And the commitment becomes structural: a few hundred seats is a line item while several thousand is a material commitment that shows up in board reporting, which changes who is watching and what happens if adoption disappoints. That third one cuts both ways, and the way to survive it is to have published the expected curve in advance rather than explaining it afterwards.

The ramp, in tranches 3:13

The ramp in tranches, and what each tranche needs before it releases. The population: named roles from the persona model rather than a headcount, because named people can be followed up and a number cannot. The threshold: measured adoption in the previous tranche, where thirty percent active is the hinge from session twenty. The window: how long the tranche runs before assessment, and six months is the right length because active use settles at twenty five to forty five percent by then. The rate: held for full population volume, in writing, which is the session nineteen structure of taking the rate and staging the volume. And the exit: what happens to seats that miss the threshold, agreed before release, because agreeing it afterwards is a request rather than a term. That last row is the one most ramps omit, and omitting it turns a staged rollout into a slow full rollout.

Knowledge check 1 4:20

First check. Tranche two hits twenty two percent active against a thirty percent threshold. What happens? A, release tranche three anyway, adoption is still building. B, apply the exit that was agreed before release: hold the next tranche, investigate why this population differs, and release or reclaim these seats on the agreed rule. C, cancel the whole programme. D, lower the threshold to twenty percent. Pause it, and as you think, ask yourself what the threshold was actually for, if it can be moved the moment it is missed.

The answer is B, and D is the one to name explicitly because it is the most common outcome in practice. Moving a threshold when it is missed converts the entire ramp into theatre: every subsequent tranche will release regardless of evidence, and everybody involved will know it, including the account team. A is the same thing with less honesty attached. C overcorrects on a single tranche when the likelier explanation is a population mismatch rather than a product failure, and twenty two percent means roughly a fifth of those people are getting real value. The useful work sits inside B, which is to investigate why this population differs from the first one. Usually it is a role mix problem, sometimes a deployment or training gap, and occasionally the population was chosen for enthusiasm rather than for evidence. Then apply the rule you wrote down, which is what makes the next threshold mean anything at all.

The consumption tail 8:31

The consumption tail, five things about the bill that arrives second. It starts after the rollout, because agents get built once people are comfortable with the product, which means the metered layer grows on a different curve from the seats and about a year behind them. It has no headcount ceiling, which is session eighteen's arithmetic at estate scale, where a single scheduled agent over two thousand records annualised to around three hundred and sixty thousand dollars before discount. The allowance goes early, since agentic runs consume five to ten times an interactive prompt and allowances ran out one to two quarters ahead of forecast. It is nobody's line, because the seats belong to procurement and the agents belong to whoever built them, so the tail grows in the gap between two owners. And it is capped by terms rather than by policy: ceiling, rollover, alert threshold, worth twenty to thirty five percent.

Guest analyst: year one, and the renewal that followed 7:02

Guest analyst  I sat in on a first Copilot renewal last year that I thought was as close to a model outcome as I have seen, and none of it was clever. A bank, about nine thousand Copilot seats deployed over four tranches across fourteen months. What they brought to that renewal was three things. A chart of monthly active users by persona for every one of those fourteen months, so you could see exactly where adoption rose and where it flattened. A cost per active user figure that they had computed themselves, full stack, which came out at a little over twice the headline. And a record of the tranche thresholds, including the one tranche that had missed and the four hundred seats they had reclaimed as a result, which they had done quietly at the time rather than arguing about it. Now the account team arrived with an expansion proposal for the remaining estate, which is exactly what you would expect. And the response was not a negotiation in the usual sense. The licence manager put her persona chart on the screen and said, these four roles sustained use above forty percent and we would like to expand there, these three sat under fifteen and we are not expanding into them, and here is the count that follows. There was very little to argue with, because it was all their data. They expanded by about a third of what was proposed, at a better rate than the original deal, and the whole conversation took two meetings.

The consumption tail 8:31

Fourteen months of persona level data, one reclaimed tranche, and a two meeting renewal. That is what year one buys you. Second check.

Knowledge check 2 8:44

Check two. Your Copilot seats plateaued six months ago but the total bill is still climbing. What is happening? A, a pricing change must have been applied. B, the consumption tail: agents and metered work built after the rollout, growing on a different curve from the seats and with no headcount ceiling. C, more users were assigned without approval. D, nothing, some month to month variance is normal. Pause it. The seats are flat, so ask yourself what else in this stack is capable of growing on its own.

The answer is B, and this is the session twenty four pattern arriving in the AI stack: cost generated by people who are not buying anything. Once a population is comfortable with Copilot, somebody automates something, and the metered layer starts growing on its own curve without passing through a purchasing decision. A is worth ruling out in about a minute and rarely explains a sustained climb. C is checkable directly from the assignment report and would show as a seat count change, which the question has already excluded. D is the answer that lets it run for another three quarters, and the difference between variance and a trend is precisely what a monthly series tells you and a single month does not. The response is the session thirty four control set: attribute the metered line, review it monthly beside the seats, and run the session eighteen arithmetic on anything scheduled.

The renewal after year one 10:23

The renewal after year one, three things that decide it, and this renewal is genuinely different from the first purchase because now there is evidence, or there is not. Your telemetry or their narrative: estates arriving with documented weekly active usage negotiated the seat estate from their own numbers, while estates without measurement negotiated against an adoption narrative that is generous about the future and silent about the idle thirty to forty five percent. The pooled right sized count: not last year's number and not the vendor's projection, but the count your usage data supports with the roles named, which is the session twenty eight counter quote applied to one product. And the tail on the table: the metered line now has a year of actuals behind it, so rate caps, rollover, and a ceiling can be negotiated against real consumption rather than against a forecast, which is a much stronger position.

What you take into it 11:25

What you take into it, the evidence file one year on. Monthly active by persona, proving which roles sustained use and at what rate, from session twenty and gathered monthly from day one. Cost per active user, proving your real price at one point eight to three point four times headline, from session thirty four's full stack division. Tranche outcomes, proving which thresholds were met and which were not, from this session's ramp, if the exits were actually applied. Consumption actuals, a year of metered spend by team, from sessions seventeen and thirty four, attributed monthly. And the retirement record, anything Copilot replaced with dates against it, which is session twelve's discipline applied to AI. Every row says the same thing, which is gathered as you went, because a file assembled in the six weeks before a renewal is a reconstruction and it reads as one to anybody experienced.

Knowledge check 3 12:28

Last check. At renewal, the account team proposes expanding Copilot to the remaining estate. What decides your answer? A, whether the discount on expansion is attractive. B, whether the roles in the remaining estate resemble the roles that sustained use in year one, because adoption follows the work rather than the offer. C, whether the budget has room. D, whether competitors have deployed at that scale. Pause it. You now have a full year of evidence about which of your own roles sustained use, so the question is whether you are going to use it.

The answer is B. After a year you know something you did not know before, which is which of your own roles sustained use, and that is a far better predictor than anything else available. If the remaining estate is mostly operations, field, and frontline work, the reliable return is not there, and no discount changes that, because the value concentrates in heavy document and email roles per session twenty. A is the standard framing and it prices the wrong question entirely, since a discount on capability nobody uses is still spending. C confuses affordability with justification, which is precisely how a thirty to forty five percent idle rate gets funded year after year. D is peer pressure dressed up as strategy, and it is silent on whether those deployments actually worked, because organisations do not publish their adoption rates. Compare the remaining roles against the roles that worked, expand where they match, and say plainly where they do not.

The scale method 14:18

The scale method, three disciplines, and none of them are new, they are module four held for twelve months rather than twelve weeks. One, publish the expected curve: before the first tranche, write down what adoption you expect and when, using the twenty five to forty five percent band as the honest reference, because a curve published in advance turns a disappointing month into a data point rather than an incident. Two, run the tranches properly: named populations, measured thresholds, agreed exits, rate held for full volume, and apply the exit when a threshold is missed, because the first time you do not, the ramp stops meaning anything to anybody. Three, watch the tail from day one: the metered line in the same monthly review as the seats, with attribution, a ceiling, rollover, and an alert. And publish the roles that did not benefit, because it costs nothing and it is the most credible thing in the paper.

Recap 15:20

Session thirty eight, three sentences. One: scale changes the failure mode from a wasted pilot to a locked commitment against a non uniform adoption curve, and the defence is a published expected curve plus tranches with measured thresholds and agreed exits. Two: the consumption tail grows after the seats plateau, has no headcount ceiling, and is capped by terms rather than by policy, so a ceiling, rollover, and alert threshold in the order form are worth twenty to thirty five percent of effective overage cost. Three: the first full year decides the renewal, because estates with documented weekly active usage negotiated from their own telemetry while estates without it negotiated against an adoption narrative that is silent about the idle thirty to forty five percent. Two sessions left. Next is what happens when the company itself changes shape.

Homework 16:24

Homework, about an hour, and this week you check the shape of your deployment. One, split adoption by persona rather than reporting the estate average, because there are a dozen curves inside that one number and the spread is the finding. Two, plot seats against the metered line, twelve months of each on one chart, and if the second is rising while the first is flat then you have found your tail. Three, check the exits: was a rule ever agreed for tranches that miss their threshold, and has it ever actually been applied, because both answers tell you something. Four, compute cost per active user, full stack over sustained weekly actives, compared against the one point eight to three point four band. Five, list the roles that did not benefit, and write them down now rather than at the renewal, because that list is what makes the rest of your paper believable.

Further reading 17:24

Five reads before next session, all free on redress compliance dot com. First, the Copilot true cost analysis, for the realised multiple and the idle rate at deployment scale. Second, the Copilot rollout bank case study, which is a large deployment with the ramp and the outcomes written up, and it is close to the story in this session. Third, the Agent 365 licensing guide, for the tail and the arithmetic that every scheduled agent needs before it ships. Fourth, capping the AI consumption overage cliff, on the three terms that bound the metered line. And fifth, negotiating Microsoft generative AI contracts, for the renewal after year one and the structure to hold in it. Next session is corporate events and special cases: mergers, divestitures, multinational structures, education and public sector agreements, and what actually transfers when the company itself changes shape. See you there.

Learning the playbook and want it applied to your numbers? We work on contingency: 25% of what we save you. Nothing saved, nothing paid.
Review my deal