HomeTraining AcademyMicrosoft Agreements and CopilotSession 17
Microsoft Agreements and Copilot · Module 4 ยท Copilot and the AI stack · Session 17 of 40 · 18:59

Copilot Chat, agents, and consumption

The metered layer beside the subscription: credits, pay as you go, and how consumption pricing behaves once agents are in production. Three knowledge checks along the way, and 1 clip from a senior cloud advisor.

The presenter in this session is an AI generated avatar. The curriculum and guidance are real, produced by Redress Compliance analysts from our consulting engagements and market network.

What you will be able to do after this session

  • 1Two layers. The per seat subscription you priced last session, and a metered layer beside it where agentic work bills by consumption. A seat only budget covers one of them.
  • 2The three routes. Pay as you go at one cent per credit, Capacity Packs at 25,000 credits for 200 dollars per tenant per month, and a Pre-Purchase Plan discounted 5 to 20 percent.
  • 3What a task costs. Light tasks run 70 to 200 credits, medium 400 to 600, heavy over 1,500. That is roughly 70 cents to over 15 dollars, and the spread is more than seven to one.
  • 4Where forecasts break. Models that used the heavy midpoint rather than the high end understated the budget by 20 to 40 percent, and heavy tasks dominate the total.
  • 5How to size a commit. Run pay as you go for a quarter or two first, then size to the floor of your proven range, because overage simply bills at pay as you go and undersizing is the safer error.

How the session works

This is a taught session, not a talking head. The instructor works through analyst grade slides, and three times the video stops on a question with four options on screen. Pause, commit to an answer, and the next slide explains which option is right and why each of the others is wrong. Once in the session the frame splits and a senior cloud advisor gives the view from inside real Oracle negotiations, and the instructor picks the clip apart when the slides return.

Homework before the next session, about an hour

  • 1Find out what you are on. Pay as you go, packs, or a prepay pool. If it is a pool, find out how much of the current one has been consumed and how much time is left.
  • 2Estimate the task mix. For two personas, how many light, medium, and heavy tasks a month. Rough is fine. Being wrong in writing beats being vague.
  • 3Price it at the high end. Heavy tier at its top, not its midpoint, converted at one cent per credit. Note the number the midpoint would have given you.
  • 4Check the attribution. Can you see consumption by department today? If not, find out what it would take, before the number gets big enough to matter.
  • 5Set one alert. A spending threshold with a name attached to it. The cheapest control in this entire course, and the one most often missing.

Session transcript

The full narration of this session, section by section, for reading and reference. Guest analyst clips are marked.

Welcome and objectives 0:02

Welcome back, session seventeen of forty. Last session we priced the Copilot subscription and built a tranche plan around adoption. Today we deal with the half of the bill that does not arrive per seat. Because underneath the subscription there is a metered layer, where agentic work bills by consumption, and it behaves nothing like the licensing this course has covered so far. I want to be direct about why this session matters more than its position in the running order suggests. Everything up to now has had a ceiling. A seat licence has a worst case, and the worst case is everybody in the organisation. Consumption has no ceiling at all. A single scheduled process can consume more than a department of people, and it will do it quietly, on a timer, without anybody experiencing the moment as a purchase. That is a genuinely different governance problem and it needs different habits.

Five takeaways. One, two layers: the per seat subscription you priced last session, and a metered layer beside it, and a seat only budget covers one of them. Two, the three routes: pay as you go at one cent per credit, Capacity Packs at twenty five thousand credits for two hundred dollars per tenant per month, and a Pre-Purchase Plan discounted five to twenty percent. Three, what a task costs: light tasks seventy to two hundred credits, medium four hundred to six hundred, heavy over fifteen hundred, which is roughly seventy cents to over fifteen dollars, a spread of more than seven to one. Four, where forecasts break: models using the heavy midpoint rather than the high end understated budgets by twenty to forty percent. Five, how to size a commit: run pay as you go for a quarter or two first, then size to the floor of your proven range, because overage simply bills at pay as you go and undersizing is the safer error.

Two layers, one bill 2:11

Two layers, one bill. The per seat layer is predictable, negotiable, and bounded by headcount, so you know the worst case on the day you sign, because the worst case is everybody. That is the layer procurement is built to handle and it is the one everybody budgets. The metered layer is agentic work billed by consumption, pooled at the tenant, denominated in credits at one cent each, and there is no headcount ceiling anywhere in it. A single scheduled agent can consume more than a department of people, which is not an intuition that seat based licensing prepares anybody for. And here is why the difference matters, which is the sentence I would underline: seat spending is a decision made once a year, while consumption spending is thousands of small decisions made by people and by software every day, none of which feel like purchasing at the moment they happen. That is the governance problem in one line, and it is why the controls have to be structural rather than cultural.

The three ways to buy credits 3:13

Three ways to buy credits, and the same credit sits behind all of them, so this is a choice about risk rather than about product. Pay as you go: one cent per credit, billed monthly in arrears for exactly what you consumed, no up front purchase, no expiry, and the risk sits with Microsoft because they are paid only for real usage. Capacity Packs: twenty five thousand credits for two hundred dollars per tenant per month, bought in the admin centre, and credits reset monthly rather than rolling over, so packs suit steady predictable volume you would rather see as a fixed line. Pre-Purchase Plan: an annual pool bought up front at a five to twenty percent discount, with unused credits expiring at term end, and there the risk moves squarely to you, because the discount is paid for with lockup. Overage bills at pay as you go rates whatever route you chose, which is the asymmetry that decides everything. And pay as you go decrements your Azure commitment, which makes departmental attribution possible.

Knowledge check 1 4:20

First check. You are offered a twenty percent discount on an annual credit pool sized to the vendor's forecast. What is the risk? A, none, twenty percent is the maximum discount available. B, forecast based pools expired with ten to thirty percent unused in the renewals we worked, which wipes out the headline discount and then some. C, the risk is that you exceed the pool. D, the risk is only that credits might get cheaper later. Pause it, and while you think, put the discount you gain beside the share of the pool that typically expires, because those two numbers are directly comparable.

The answer is B, and the arithmetic is not close. The discount tops out at twenty percent. Forecast sized pools expired with ten to thirty percent unused across the commitment renewals behind this course. At the upper end of that range you have paid a premium for the privilege of committing early, and the discount was the entire reason you committed. Now C is the risk everybody worries about, and it is the wrong one, which is the most useful thing in this slide. Overage simply bills at pay as you go rates. Exceeding the pool costs you the standard price and nothing more. So the asymmetry is: undershooting a pool is cheap, overshooting it is expensive, and when you are uncertain you commit low. A treats the maximum available discount as though it were free money. D is a real consideration and a minor one beside a third of your pool evaporating at term end without anybody sending you a warning.

What a task actually costs 8:38

What a task actually costs, five things about the meter. It bills by difficulty, not by what you sent, so the unit to think in is the task rather than the prompt. The published ranges: light tasks seventy to two hundred credits, medium four hundred to six hundred, heavy over fifteen hundred, which at one cent per credit is roughly seventy cents to over fifteen dollars per task. The heavy tier decides everything, because the spread is more than seven to one, and in the deployments we modelled a small number of heavy analyses by power users drove the majority of total credit spend. Model the high end rather than the midpoint: models using the heavy midpoint understated budgets by twenty to forty percent, and if you take one modelling rule from this session take that one. And API calls are a separate line, because Work IQ API calls bill per agent beside the task credits. The method that follows: forecast by persona rather than by headcount.

Guest analyst: the pool that expired 7:10

Guest analyst  The credit pool conversation I always come back to happened at an engineering group about a year into their Copilot deployment. They had bought an annual pre purchase pool, taken the full discount, and the pool had been sized from a forecast the account team had built with them. Not imposed on them, built with them, which is worth saying because everyone was acting in good faith. Ten months in, their finance business partner asked a simple question: how much of the pool have we used. Nobody knew, and it took about a week to find out, which was itself the finding. The answer was sixty seven percent, with two months left, and no realistic path to burning the rest. So they had taken a discount in the high teens and they were going to expire about a third of the pool. Net, they had paid more per credit than pay as you go would have cost them, for the privilege of committing a year early. What I found interesting was the reason the forecast was wrong. It was not that usage was low. Their light and medium task volume had actually come in above forecast. What had not materialised was a set of heavy analytical workloads that three teams had said they would run, and those were where most of the forecast credits lived. The lesson they took, and I think it is the right one, was not to distrust the vendor. It was that a pool should be sized to the work you are already doing, not to the work somebody intends to start.

What a task actually costs 8:38

Sixty seven percent burned, a third expired, and a net price above pay as you go. Size to what you already do. Second check.

Knowledge check 2 8:50

Check two. Your forecast assumes the average task. Actual usage turns out to be eighty percent light tasks and five percent heavy. Is the average safe? A, yes, five percent heavy is a small share of activity. B, no, because the spread is more than seven to one, so that five percent can dominate the total, and the budget is decided by the heavy tier rather than by the typical task. C, yes, provided the average is calculated correctly. D, no, but the effect is small enough to absorb. Pause it, and before you answer, actually multiply the share of tasks by the credits each tier consumes, because the answer is more extreme than it looks.

The answer is B. Do the multiplication. A heavy task at fifteen hundred credits or more against a light task at seventy to two hundred is a ratio well beyond seven to one, so a small population of heavy analyses can carry more total credit than the entire large population of light ones. In the deployments modelled for this course that is exactly what happened: a small number of heavy analyses by power users drove the majority of total spend. A and C both treat frequency as though it were cost, which is the averaging error from session fifteen arriving in a completely new place, and it is why a task count tells you almost nothing about a bill. D is directionally right and quantitatively wrong, and I want to be precise about that: understating the heavy tier moved budgets by twenty to forty percent, which is not an absorbable variance on a line that has no headcount ceiling underneath it to catch you.

The forecasting traps 10:41

Three forecasting traps, all avoidable and all common. Averaging the heavy task down: the biggest one. Budgets that held forecast the heavy tier honestly at its high end, and budgets that blew through cap had averaged it down, usually in the name of not being alarmist, and they were understated by twenty to forty percent. Sizing on the vendor forecast: prepay pools sized that way were the single most common source of waste across the commitment renewals we worked, and the forecast is not dishonest, it is simply produced by a party whose interest is a larger pool. And forgetting the scheduled work: a person consumes credits when they choose to, while a scheduled agent consumes them on a timer, forever, whether anybody is watching or not, so any recurring automation needs its own forecast line. Notice the pattern across all three. Consumption pricing punishes optimism in a way seat licensing does not, because a seat you overbuy costs you that seat while a wrong consumption assumption compounds daily.

Sizing a commitment 11:51

Sizing a commitment, the method in order. One, run pay as you go for one or two quarters of real consumption first, because buyers who did this sized their eventual commit far more accurately than those who did not. Two, build the persona model: task mix per persona with the heavy tier at its high end, which gives you a forecast you can explain and correct rather than defend. Three, find the floor: the lowest monthly consumption in your proven range, because the floor is what you can commit to safely. Four, commit to the floor, never to the forecast, since overage bills at pay as you go and therefore undersizing costs you nothing extra. Five, validate quarterly, actuals against model, adjusting the persona assumptions as you learn. Step four is the one that gets argued with, because committing to the floor feels like leaving a discount on the table. It is not. It is buying the discount only on volume you are certain of, and paying standard rates on volume you are not.

Knowledge check 3 12:59

Last check. Consumption is new in your tenant and finance wants a fixed annual number. What do you recommend? A, commit to a prepay pool for the whole forecast so the number is fixed. B, pay as you go for a quarter or two to establish the real range, then a commitment sized to the proven floor, with the variable remainder budgeted explicitly. C, Capacity Packs for everything, since they are a fixed monthly line. D, refuse to deploy agents until pricing is predictable. Pause it. Finance wants predictability and that is a completely reasonable thing to want, so the question is what you can honestly make predictable and what you cannot.

The answer is B. Finance's request is legitimate, and the honest answer is that consumption has a predictable floor and a variable remainder, so you fix the part that is fixable and make the rest visible rather than pretending it away. A produces a fixed number by converting forecast risk into expired credits, which is how pools finish the year ten to thirty percent unused, and the fixed number was never really fixed, it was just paid in advance. C is worth considering rather than dismissing, because packs genuinely do suit steady predictable monthly volume, but they reset monthly rather than rolling over, so buying packs for everything before you know your pattern just repeats the same overcommitment in smaller instalments. D is not available in most organisations and would not be right anyway, because pay as you go with a spending alert is a perfectly responsible way to begin: it is the one route where you cannot commit to something you never use.

The consumption method 14:46

The consumption method, three habits, and they are different from everything else in this course because the bill is generated continuously rather than annually. Meter before you commit: pay as you go first, always, for at least a quarter, because it is the only route that produces the data every other route depends on, and the premium you pay for that quarter is small against a mis sized annual pool. Attribute from day one: pay as you go decrements the Azure commitment and Azure meters the spend, so departmental attribution is available if you set it up at the start, and attribution changes behaviour in a way that a central budget line never does. And review the heavy tail monthly: the heavy tasks are the budget, so a monthly look at the largest consumers, by person and by agent, catches the pattern change that turns a stable line into an overrun, usually a month or two before the invoice does. Govern this like cloud spend, because that is what it is.

Recap 15:52

Session seventeen, three sentences. One: there are two layers to the Copilot bill, a per seat subscription bounded by headcount and a metered layer with no ceiling at all, and a budget built from seats alone covers only one of them. Two: credits bill by task difficulty at one cent each, light seventy to two hundred, medium four hundred to six hundred, heavy over fifteen hundred, and because the spread exceeds seven to one the heavy tier decides the budget and must be modelled at its high end. Three: run pay as you go for a quarter or two, size any commitment to the floor of the proven range rather than the vendor forecast, and let overage bill at standard rates, because undersizing is the cheaper error. Next session takes the metered layer to its logical conclusion, which is what happens when the thing consuming the software is not a person at all.

Homework 16:54

Homework, about an hour, and this week you build the consumption model. One, find out what you are on: pay as you go, packs, or a prepay pool, and if it is a pool find out how much of the current one has been consumed and how much time is left, because that question took a week to answer at the engineering group in the story. Two, estimate the task mix for two personas, how many light, medium, and heavy tasks a month, and rough is fine because being wrong in writing beats being vague. Three, price it at the high end, heavy tier at its top rather than its midpoint, converted at one cent per credit, and note what the midpoint would have given you so you can see the size of that choice. Four, check the attribution: can you see consumption by department today, and if not, what would it take. Five, set one spending alert with a name attached to it. Cheapest control in the course.

Further reading 17:52

Five reads before next session, all free on redress compliance dot com. First, Copilot Credits, pay as you go against pre purchase, which covers the three routes, when prepay genuinely pays, and how to size a commit you will actually burn. Second, Copilot Credits cost per task, for the task tiers, the modelling method, and why the heavy tier decides the budget. Third, Copilot Credits and the Azure commitment, on how credit spend interacts with your Azure commitment, which is session twenty one. Fourth, Copilot Credits governance at the EA, for the contract terms that make the metered layer manageable rather than merely observable. And fifth, capping the AI consumption overage cliff, on the controls that stop a metered line becoming an incident. Next session is Copilot cowork: agents working alongside people in the flow of work, and what changes about the licence question when the thing consuming the software is not a person. See you there.

Learning the playbook and want it applied to your numbers? We work on contingency: 25% of what we save you. Nothing saved, nothing paid.
Review my deal