What draws on the pool and how fast, why seat forecasts understate it by construction, and the four terms that cap the variable bill. Three knowledge checks along the way, and 3 clips from a senior licensing analyst.
This is a taught session, not a talking head. The instructor works through analyst grade slides, and three times the video stops on a question with four options on screen. Pause, commit to an answer, and the next slide explains which option is right and why each of the others is wrong. 3 times in the session the frame splits and a senior licensing analyst gives the view from inside real ServiceNow negotiations, and the instructor picks the clip apart when the slides return.
The full narration of this session, section by section, for reading and reference. Guest analyst clips are marked.
Welcome back, session eight, and today we take apart the meter. I have been promising this session since session one, when I told you ServiceNow now prices three things at once and the third one, consumption, is the one that moves. Everything since has circled it. Today we land on it properly. What actually draws on the assist pool and how fast, why the forecast almost everyone builds is wrong in a specific and predictable direction, where credits vanish without producing anything, and the four contract terms that put a ceiling on the variable half of your bill. I will give you the sentence that organizes the whole session right now, and it comes from the buyers who learned it expensively. The pool is a budget, not an entitlement. It is a finite quantity that gets spent, and when it is gone the meter keeps running and the invoice keeps counting. Three checks, homework that builds you a real consumption model, and by the end you will be able to price an AI line properly. Let's start.
Five objectives. First, read the pool correctly, understanding why the bundled allocation is a budget rather than an entitlement, and exactly what happens the moment it is spent. Second, rank the burn, ordering the four things that draw on your pool by how fast they consume it, and identifying which ones nobody models, because that gap is where the surprises live. Third, build an actions model, a forecast based on prompts, agent tasks, and non production load instead of seat count, which takes an afternoon and changes every number in your negotiation. Fourth, find the leakage, the places credits disappear without producing any value at all, and what fixes each one. And fifth, land the four terms, pool sized to a modeled year, a capped unit rate, rollover, and an annual ceiling. By the end of today the AI line stops being a mystery on your quote and becomes a number you can defend to your CFO.
Four numbers. Twenty two percent. That is the median overage against tier spend once Now Assist and agents were genuinely in use, inside a range of fifteen to thirty percent. Put that plainly, the average estate ended up paying about a fifth again on top of a tier price it had negotiated carefully, on a meter it had not negotiated at all. Thirty five to fifty percent. That is how far credit pool commitments overshot actual first year draw. Now hold those two numbers together, because they look contradictory and they are not. Some estates burned far more than they expected and some committed far more than they used, and both are the same failure, forecasting without a model. Three to five x. The variation in credits consumed per assisted action across similar workflows, meaning the identical task costs several times more in one estate than in another, and we will get to why. And twenty to forty percent, what buyer side governance saved on rollout cost. The note underneath matters. Under buy and you meet overage at the true up. Over buy and a third of estates watched prepaid packs simply expire. Only a model puts you between them. Let's hear how that plays out.
Guest analyst clip. The phrase I have started using with clients is that the pool is a budget wearing the costume of an entitlement. And the costume matters, because of where it sits on the paper. It arrives bundled in your tier, described in the same breath as features you own outright, and human beings read anything bundled as included. So the platform team hears, AI is included, go and use it, and they do exactly that, enthusiastically, which is what you wanted. Then somewhere around month five the finance business partner asks why there is a variable line on a SaaS invoice, and nobody in the room can explain the number because nobody was watching a meter they did not know was running. I sat in one of those meetings where the platform owner genuinely believed the assists were unlimited within the tier. He was not careless, he had read the material, and the material really does present it as included. Nobody had told him there was a quantity. And here is the part that should worry a buyer most. That customer had negotiated a very good tier price. Their per user rate was better than most estates I benchmark. It made no difference at all, because the money moved to the line they never looked at. You can win the visible negotiation completely and still lose the deal on the meter.
You can win the visible negotiation completely and still lose the deal on the meter. And notice the mechanism in that story, nobody was careless. The platform owner read the material and the material genuinely reads as included. That is why this session exists as a session rather than a footnote. The word bundled does real psychological work, and it is doing it inside your organisation right now, in the head of whoever runs your platform. Now, what actually burns the pool.
Four sources, ranked, and watch the right hand column because it is the whole lesson. Interactive prompt, low burn, that is a summary, a suggestion, a drafted reply, a person clicking a button and getting help. Usually modeled, and usually the only thing modeled. Agentic workflow task, medium to high burn, because one task chains multiple actions inside itself. Rarely modeled. Autonomous agent run, high burn, several times an interactive prompt, many actions per completed task. Rarely modeled. And development and sub production, variable and, this is the important word, unbounded by production adoption. A test harness runs harder than any human user, tirelessly, at whatever schedule somebody set. Almost never modeled. Look at that column again. The two heaviest consumers on the list are the two nobody forecasts, and the fourth one is not even bounded by anything happening in your live business. The principle underneath is one sentence. The pool depletes by actions, not by people. An agent on a schedule runs whether anyone is watching, on a Sunday, over Christmas, at three in the morning.
So why do seat based forecasts fail? Not by accident, and not by a little. They fail by construction, five reasons. Seats do not consume, actions do, and a forecast built on licensed users assumes consumption scales with headcount when it actually scales with automation, which is deliberately decoupled from headcount, that is the entire point of automation. Agents break the ratio, one agent completing one task can consume several times what a person consumes, and it runs on a schedule rather than a working day, so it has no evenings and no holidays. Non production is invisible, dev and sub production draw the production meter, and heavy agent testing is the single most common source of unexpected overage, teams have exhausted pools without one production user noticing anything. Adoption misleads in both directions, and this is subtle, fulfiller adoption ran thirty to fifty five percent below the SKU description on broad rollouts, so when somebody tells you adoption is low and therefore the pool is safe, they have it backwards, low adoption alongside high burn means your burn per action is terrible. And finally, the seller forecasts seats too. The consumption assumption behind your quote is almost always a per user figure. Ask what action volume that implies. The conversation changes immediately, because usually nobody has done that arithmetic. First check.
Knowledge check one. Six weeks after go live, your pool is sixty percent consumed, but only twenty percent of licensed users have touched Now Assist at all. What is the most likely explanation? A, the counters are wrong, because low adoption cannot produce that burn. B, agent runs and non production testing are drawing on the same pool. C, the twenty percent are power users generating enormous prompt volume. Or D, the pool was mis provisioned and ServiceNow will credit it back. Pause here, and ask what draws on the pool without a person present.
The answer is B, and this is the signature pattern of the whole model, burn decoupled from adoption. Scheduled agents run without users, and dev and sub production draw the production meter, so a test harness in a non production instance can comfortably outrun your entire live population. Answer C is arithmetically possible but rare, interactive prompts sit at the low end of the burn range and you would need genuinely extraordinary volume. Answer A is session five's losing argument returning in a new costume, disputing counters that read your own instance. And answer D assumes goodwill where the contract says usage is usage, and the usage was real, it just was not human. So the practical rule, when the pool moves faster than adoption explains, check the environment split first. Every time. It is the first question, not the last.
The actions model, four inputs, one number, and this genuinely takes an afternoon. Input one, prompts per user, interactive assists per active user per month, and take that from your own pilot data rather than the SKU description, then multiply by realistic adoption rather than by licensed seats. Input two, agent tasks times the multiplier, planned agent task volume multiplied by the actions each task chains, and once agents reach production this line dominates the model, often by an order of magnitude. Input three, the non production load, dev, test, sub production, including the load tests, and I want to be blunt about why it goes in the model. It is in the invoice either way. Excluding it does not make it free, it only makes your forecast wrong. And input four, the ramp, adoption over time rather than at steady state, because a pool sized for month twelve is wasted in month one, and a pool sized for month one fails you by month six. Then compare the result against the bundled pool for each tier. That comparison, done before signature, is the entire content of a consumption negotiation. Everything else is decoration. Second check.
Knowledge check two. Your model says a realistic year burns roughly twice the bundled pool. Which move is the worst one? A, negotiate a larger pool at signature. B, cap the overage unit rate for the term. C, accept the standard pool and pay overage as it arrives. Or D, buy a bigger pool and secure rollover for whatever you do not use. Pause, and ask which option leaves you with a price you cannot control.
The answer is C, accepting the standard pool and paying as you go, and it is worth being precise about why it is the worst rather than merely suboptimal. You have already forecast that the meter will run hot. Choosing to pay overage anyway hands the vendor an uncapped variable line on consumption you predicted, and overage rates are negotiable at signature and effectively impossible to negotiate later, because by then your usage is live, visible to them, and rising. You have removed your own leverage with information you already had. A, B, and D are all defensible, and they are best combined, with D the strongest because rollover protects against the over buy risk, the stranded prepaid packs. Here is the general principle. Knowing the burn and doing nothing contractual about it is the one genuinely indefensible position, because ignorance at least explains itself.
Now leakage, where credits disappear without producing value, five sources. Weak knowledge bases, and this is the big one, knowledge quality drove sixty to eighty percent of the variance in per skill performance. Think about what that means, most of the three to five fold difference in cost per action between estates is not the AI, it is the quality of what you fed it. Thin inputs generate long completions that help nobody and burn pool doing it. Routing loops, assists invoked repeatedly on work bouncing between queues, so you pay several times for one unresolved problem. Ungoverned agent scope, agents pointed at high volume low value work because it was easy to automate rather than worth automating, which is a governance failure dressed as a technology success. Test harnesses nobody throttled, the fastest way to empty a pool and the least likely to be noticed until the invoice. And broad rollout without a pilot, estate wide launches that skipped validation rarely showed credible per fulfiller value before the first renewal, while a scoped pilot, fifty to a hundred fifty fulfillers across two or three product areas, proved the unit economics in ninety days. Notice that four of those five are fixable by you, internally, with no vendor conversation at all.
Guest analyst clip. The knowledge base finding is the one I wish more executives understood, because it inverts how people think about AI cost. Everybody assumes the price of AI is set by the vendor. In consumption models a surprisingly large share of it is set by you, specifically by the quality of what you point the model at. Here is the pattern we measured. Two customers, same product, same skills enabled, similar workflows. One had mature knowledge management, articles that were current, deduplicated, actually written for resolution. The other had a knowledge base that had not been curated in years, contradictory articles, dead links, half of it out of date. The second customer's cost per resolved interaction ran several times higher, and the reason is almost mechanical. Weak inputs produce long, hedging, uncertain completions, and those completions consume more. Then, because the answer was unhelpful, the human tries again, which consumes again. So you pay twice for the same failure, once for the bad answer and once for the retry. And the recommendation lands strangely with leadership, because the highest return AI cost control we recommend is usually not an AI project at all. It is a knowledge remediation project. Clean the inputs and the meter slows down on its own, without touching a single contract term.
The highest return AI cost control is usually a knowledge remediation project. Sit with that, because it reframes the ownership question. If a meaningful share of your AI spend is set by the quality of your own content, then the AI budget is partly a knowledge management budget, and those two things usually live with different people who do not talk. Whoever owns your ServiceNow licensing needs a line into whoever owns knowledge. Now, the four terms.
Four terms that cap the variable half of the bill. Pool sized to a modeled year, meaning the allocation is set against your actions model rather than the tier default, and the argument for it is simple, headroom bought at signature is cheaper than overage bought under pressure. A capped overage unit rate, fixing the per unit price for the term so the variable line has a known ceiling, and remember the asymmetry, negotiable before signature, effectively impossible once the meter is running hot. Rollover, unused capacity carrying forward instead of expiring, which is the term that faces the over buy risk, and given that prepaid packs expired unused in about a third of estates, it is worth more than it sounds. And an annual spend ceiling, a threshold that triggers a review conversation rather than an automatic invoice, and I like this one particularly because it converts a financial surprise into a scheduled decision, and it gives your internal governance a contractual hook to point at. Now the note under the table, and this is the single most actionable fact in the session. Credit carry forward was available on roughly half the contracts we ran when it was asked for explicitly. It is not in the default template. It was worth twelve to twenty two percent across the term on estates with uneven adoption. Half. Just for asking. Third check.
Knowledge check three, and it deliberately points at the failure nobody plans for. Adoption is slower than you planned, and you will finish the year comfortably under your pool. Which term protects you? A, the capped overage unit rate. B, rollover, so unused capacity carries forward. C, the annual spend ceiling. Or D, nothing, under use is always the customer's loss. Pause here.
The answer is B, rollover, and it is the only term on that list facing the under use direction at all. The rate cap and the ceiling both protect you against burning too much, which is the risk everybody anticipates because it is the frightening one. But under use is the more common failure, pool commitments overshot first year draw by thirty five to fifty percent, and packs expired unused in about a third of estates. So the money most often lost on this meter is money spent on capacity that quietly evaporated, not money spent on overage. Answer D is the default contractual position, and stating it that baldly shows you why the term is worth asking for. It was granted on roughly half of contracts where somebody asked explicitly, and it is never, ever offered. Both directions of the forecast cost you. Only one of them is frightening enough that people remember to negotiate it.
The clause set caps the bill. Governance decides whether you ever reach the cap, and it is where the twenty to forty percent saving actually comes from. Four practices. Pilot before estate, fifty to a hundred fifty fulfillers across two or three product areas, ninety days, measuring credits per resolved interaction, and that number, yours, not a vendor benchmark, sizes everything that follows. Per wave ceilings, each rollout wave with a credit ceiling and a named owner, and a wave that exceeds its ceiling stops for review rather than quietly borrowing from next quarter, which is what happens by default. Fix knowledge first, because knowledge quality drives most of the per skill variance, which makes knowledge remediation a cost control rather than a nice to have, and lets you fund it accordingly. And throttle non production, dev and test schedules capped and owned, exactly like the ITOM discovery scopes in session five. Same failure mode, same fix, an unowned schedule consuming quietly. Then one metric to report monthly, beside the pool trend. Credits per resolved interaction. It is the one number that tells you whether the AI is producing value or just producing invoices, and almost nobody tracks it.
Guest analyst clip. I want to close on the governance point because it is where I see the biggest gap between intention and practice. Every customer I work with agrees, in principle, that AI consumption needs governing. Very few have anybody actually doing it, and the reason is structural rather than lazy. Consumption sits in the gap between three teams. The platform team owns the technology but not the budget. Finance owns the budget but cannot read the meter. And procurement owns the contract but has usually moved on to the next renewal. So the meter runs in a space where everyone assumes somebody else is watching. The estates that control this well do one unglamorous thing, they give the number an owner, the same way session two gave the user count an owner. One person receives the consumption report monthly, credits per resolved interaction, burn by environment, pool trend against the ceiling. It takes them under an hour. And when a wave starts burning three times what the pilot predicted, they find out in week two rather than at the true up, when it is still a technical conversation about a misconfigured agent rather than a commercial conversation about an invoice. That is the whole difference. Not sophistication, ownership, and about an hour a month of somebody's attention.
The meter runs in the gap between three teams, and the fix is an owner and an hour a month. That is the same conclusion as session two's user file and session five's quarterly pass, and by now I hope the pattern is obvious. Almost nothing in this discipline requires cleverness. It requires that a specific number belongs to a specific person on a specific schedule. Let's recap.
The meter, in three sentences. The pool is a budget rather than an entitlement, and once it is spent usage bills per unit, which put overage at fifteen to thirty percent of tier spend, a twenty two percent median, once agents were live. Consumption scales with actions and automation rather than with seats, so every seat based forecast understates it by construction, and dev and sub production draw the very same meter as production. And four terms cap the variable half of the bill, a pool sized to a modeled year, a capped unit rate, rollover, and an annual ceiling, while governance wave by wave decides whether you ever reach them. Next session closes module two with the layer sitting on top of all this. AI agents. What Prime actually gates, what the AI Control Tower is for, and the part most buyers never learn, the public AI revenue target that makes your willingness to adopt worth real discount points.
Homework, about an hour, and this week you build the model. One, split the burn, take last quarter's assist consumption and split it by environment, production against dev and sub production, and note the share that never served a single user. That number surprises people. Two, build the model, prompts per active user, plus planned agent tasks times an action multiplier, plus the non production load, ramped across twelve months. Three, compare to the pool, put your modeled year against your bundled allocation and write down the gap in either direction. Four, check the four terms, read your consumption schedule for the pool sizing basis, a rate cap, rollover, and any spend ceiling, marking each present, weak, or absent. And five, find your unit economics, credits per resolved interaction for one workflow. If nobody in your organisation measures that today, then establishing it is the most valuable thing you will do this quarter, and it costs nothing but attention.
Further reading, five guides. The consumption and overage guide is today's burn table and clause set in written form, with the benchmark data behind the twenty two percent median. The credit model white paper decodes the pricing and gives you the per wave governance framework in full. The Now Assist pillar covers the adoption shape and the renewal envelope across a multi year rollout. The pricing deep dive covers the mechanics underneath the pool and how packs and tiers interact at renewal. And the insurance case study shows the pilot to estate sequence in practice, with the governance that kept the line predictable. That is session eight. Build the model, find your credits per resolved interaction, and I will see you in session nine for the agents.