Eligible technologies, partitions and shared processor pools, VMware clusters, and how one host becomes the whole estate. Three knowledge checks along the way, and 4 clips from a senior licensing analyst.
This is a taught session, not a talking head. The instructor works through analyst grade slides, and three times the video stops on a question with four options on screen. Pause, commit to an answer, and the next slide explains which option is right and why each of the others is wrong. 4 times in the session the frame splits and a senior licensing analyst gives the view from inside real IBM negotiations, and the instructor picks the clip apart when the slides return.
The full narration of this session, section by section, for reading and reference. Guest analyst clips are marked.
Welcome to session nine. Session six gave you the rule, could run rather than does run, and I said at the time that the hard part is not the rule but the question of where could run ends. That is today. Because here is what makes this session different from the two before it. The most expensive line in an IBM estate is not drawn in a contract, and it is not drawn in a licensing tool. It is drawn in a hypervisor console, by an engineer, on a Tuesday, for reasons that have nothing to do with licensing and are usually excellent. Three knowledge checks. Let's begin.
Five objectives. First, draw the countable boundary, because for any workload the set of hosts it is permitted to run on is the capacity in scope, and that set is almost always larger than the people running it expect. Second, read a cluster as a price, because a workload free to move across four sixteen core hosts is a sixty four core position, which makes cluster design a licensing decision taken by teams who are never asked about licensing. Third, use the Power levers properly, where the licensable unit is the activated core, a capped shared processor pool bounds what the partitions can consume, and the partition map rather than the chassis decides the number. Fourth, name the four ways a boundary grows, which are hosts added, mobility widened, clusters merged, and capacity activated on demand. And fifth, keep the mobility history, because it is the evidence that removes hosts from a claim.
Four numbers. Four thousand four hundred and eighty, the PVUs for one workload free to move across four sixteen core hosts, where the same workload pinned and measured is five hundred and sixty. Seven times, the full capacity difference on dense VMware clusters in one remediation, because density is what makes the ratio brutal. Eighty two to nought point six, in millions of dollars, where an opening audit claim priced every host in the cluster and the mobility history proved the software had never run on most of them. And forty percent, the share of hosts that had fallen out of coverage in that pharmaceutical estate, largely because cluster changes were never mirrored in the tool. And then the note, which is the sentence I would have you carry out of this session. Nobody reads a capacity expansion as a licensing event.
Guest analyst clip. I want to describe the moment the money gets spent, because it is not the moment anybody would guess. It is not a contract signing and it is not an audit. It is an engineer, at a console, adding a host to a cluster, on a Tuesday afternoon, because a monitoring alert said memory pressure. That is the moment. Everything downstream of it, the finding, the claim, the settlement, is just the consequence arriving late. Now, what strikes me about that engineer is that they did their job correctly. They were asked to keep a service healthy and they kept it healthy, using the standard change process, with an approval from their own management chain. Nobody in that process was competent to ask the licensing question, and nobody in the licensing function was in the process at all. So the exposure was created by a well run organisation following its own rules. Which tells you the fix cannot be training people to care more, because they already care about the things they were asked to care about. The fix is structural. The licensing question has to be a field on the form that the process already uses, so that it gets asked by the workflow rather than by somebody's conscience.
Created by a well run organisation following its own rules. Which is why the fix is structural rather than cultural. Now let me show you what the boundary is actually worth.
One eight core workload, four boundaries, four prices, all at seventy PVU per core. Measured and pinned to one host, you count the eight virtual cores the workload can use, which is five hundred and sixty PVUs. Measured on one sixteen core host where the workload can reach all of it, you count the host, one thousand one hundred and twenty. Unmeasured on a sixty four core host, you count every physical core in the machine, four thousand four hundred and eighty. And free to move across four sixteen core hosts, you count every core in the cluster it can reach, which is the same four thousand four hundred and eighty. Now look at the note underneath, because that is the whole session. The workload is identical in all four rows. Nothing about it changed. What changed is where it was permitted to run, and whether anybody measured it. Those are two separate decisions, and both of them belong to somebody.
Knowledge check one. Infrastructure adds two hosts to a VMware cluster to relieve memory pressure. An IBM product runs on one VM in that cluster. What just happened? A, nothing, because the workload did not move and its resource allocation is unchanged. B, the countable boundary grew by two hosts, because the workload can now be scheduled onto them. C, nothing, provided the tool is deployed on the new hosts. D, the cluster must be re licensed from scratch at full capacity. Pause here, and ask what determines scope, the workload or the permission.
The answer is B. Where a workload can move, the capacity it could run on is the capacity in scope, so a capacity expansion is a licensing event even though nobody involved would ever describe it that way. Answer C is the interesting one, because it is half right and that makes it the more dangerous answer in the room. Putting the tool on those hosts keeps you measured, and being measured matters enormously, as the last two sessions were entirely about. But measurement does not shrink a boundary you widened. It only means you can prove how wide it is. Shrinking it takes affinity, isolation, or a smaller cluster.
So, VMware, where the unit is the cluster and not the host. Mobility defines the edge, so if a virtual machine can be migrated to a host, that host's cores can enter the calculation, and both live migration and automated placement count as can. Density makes it brutal, because the better consolidated the cluster the worse the ratio, and on dense clusters the full capacity difference reached a factor of seven in one remediation. A handful of virtual machines prices every host, which is exactly how the largest audit defence in our file reached eighty two million dollars, with products on a few VMs priced across every host in the cluster. Automated placement is the quiet part, because rules that rebalance workloads for performance are excellent engineering and they are simultaneously a standing declaration that the workload may run anywhere in the cluster. So the cluster becomes a licensing object, which means somebody has to own the question of which clusters IBM software is permitted to live in, and that question has an answer with a price on it.
Four ways the boundary grows, and notice as I go that none of them look like licensing. Hosts added for capacity, a perfectly ordinary expansion approved on infrastructure grounds that quietly enlarges the set of cores a program could run on. Mobility widened, which is an affinity rule relaxed, a placement policy simplified, a maintenance mode change, all small, all correct, all boundary changes. Clusters merged, where a consolidation project joins two clusters into one and the workload that used to reach four hosts now reaches twelve without moving an inch. Capacity activated on demand, which on Power raises the licensable footprint, and temporary activations nobody tracked become a permanent question at audit time. And each one is somebody else's decision, because not one of these is made by a licensing team, which is precisely why the control has to sit in the change process rather than in a quarterly review.
Guest analyst clip. The cluster merge is my favourite example of this whole category, because of how completely invisible it is. Two clusters, four hosts each. Somebody proposes merging them, and the business case is genuinely good: better resource utilisation, simpler management, fewer things to patch. It gets approved on those grounds, correctly. And on the day it completes, an IBM workload that could previously reach four hosts can now reach eight. It did not move. Nobody touched it. Its configuration file is byte for byte identical to what it was yesterday. And its licensable boundary has doubled. Now here is what I find genuinely difficult about advising on this, and I want to be honest about it rather than pretend the answer is easy. The merge was a good idea. I am not going to tell an infrastructure team to run two half empty clusters forever because of a licensing metric, because that is a bad trade in most estates. What I will say is that the trade should be made deliberately, with the number visible. Sometimes you merge anyway and accept the cost, because the operational saving is larger. Sometimes you merge and keep IBM workloads pinned to a subset. Both are fine. What is not fine is discovering afterwards that you made a decision you did not know you were making.
Sometimes you merge anyway and accept the cost. What is not acceptable is discovering afterwards that you made a decision you did not know you were making.
Knowledge check two. On Power, you buy a forty eight core server, activate sixteen cores, and run IBM software in a capped shared processor pool of eight. What sets the licensable number? A, the forty eight physical cores installed in the chassis. B, the configuration you run, meaning the activated cores bounded by the pool cap, provided it is measured. C, the sixteen activated cores, always, regardless of the pool. D, whatever the workload consumed on average across the quarter. Pause here, and ask which of these numbers is a decision you made and which is a fact about the metal.
The answer is B, the configuration you run. The licensable unit on Power is the activated core rather than the installed one, and a capped shared pool bounds what the partitions drawn from it can consume, so you license the cap rather than the whole server. Answer D is the consumption assumption again, and it is wrong here for exactly the same reason it was wrong in session six, because averages are not boundaries. And notice the condition sitting at the end of the correct answer, provided it is measured, because without measurement the whole arrangement falls back to physical capacity and the careful configuration work counts for nothing.
Power deserves its own treatment, because it gives you more control than VMware does. The activated core is the unit, not the installed core, so a forty eight core chassis with sixteen activated is a sixteen core licensing conversation, and that is a genuine and frequently overlooked saving. Partitions divide and pools bound, because micropartitioning draws logical partitions from a shared processor pool and a capped pool limits what those partitions can consume, which caps the exposure. Finer partitions do not lower the floor, so subdividing inside a partition, as workload partitions do on AIX, organises the workload while the licensable boundary stays the cap you set above it. Capacity on demand is the trap, since temporary activations can raise the licensable footprint and an activation nobody recorded becomes an activation nobody can explain. And the numbers are worth having, because a capped pool has run fifteen to thirty percent below full capacity, measured sub capacity twenty to forty percent below, and the two are complementary rather than alternatives.
Guest analyst clip. There is a line I use with Power estates that seems to land, so let me give it to you directly. On Power, you do not license the server you bought. You license the configuration you run, if you build it that way. And the conditional at the end is doing all the work in that sentence. Because the capability is genuinely there. You can activate a fraction of the cores. You can cap a shared pool. You can draw partitions that consume only what they need, and the paperwork will follow the configuration faithfully. What I find, though, is that estates buy the capability and then never operate it, in the sense that the initial build is thoughtful and then three years of changes are not. Cores get activated for a project and never deactivated. A pool cap gets raised during an incident, at three in the morning, entirely justifiably, and nobody lowers it afterwards because nobody owns the number. And the licensing consequence of each of those individually is small, which is why none of them get flagged, and collectively they are the difference between the configuration you designed and the configuration you are actually running. So the question I would put to any Power estate is not how did you build this. It is when did you last check that it still looks the way you built it.
The question is not how you built it, it is when you last checked that it still looks the way you built it. Which brings us to the levers themselves.
Four ways to make the boundary smaller, all available to you today without asking anybody's permission. Dedicate a cluster, because a cluster that only runs IBM software has a boundary you chose, and it is the bluntest lever and usually the largest one. Pin with affinity rules, restricting which hosts a workload may be placed on, so the unreachable hosts leave the calculation, and that is a configuration change with a price attached to it. Cap the pool and right size the cores, so on Power you cap the shared pool, and everywhere you allocate the virtual cores the workload actually needs rather than the cores that were convenient at build time. Remove what is not running, because installed counts, so a decommissioned product whose binaries remain is still a deployment, and uninstalling is the cheapest reduction available anywhere. And price each one before you choose, because every option here is calculable in advance, which makes this a design conversation with numbers in it rather than a negotiation.
Now the evidence, which is what settles arguments about boundaries in the past. The claim is a model, because an opening position prices every host the auditor cannot rule out, and that is why cluster wide counting produces such large numbers so quickly. Migration logs rule hosts out, and in the eighty two million dollar defence the vSphere migration history showed the virtual machines had never touched most of the hosts the claim priced. Partial data is still data, because that same defence found partial tool coverage had been discarded rather than supplemented, and incomplete evidence combined with other records is worth a great deal more than nothing. The first sixty days go on evidence rather than on negotiation, with vCenter inventories, tool data and change management logs rebuilding the position before anybody discussed a number. So retain the history deliberately, because migration logs, cluster membership records and change tickets all have a default retention set by an infrastructure team who do not know what they are holding.
Knowledge check three. An audit prices IBM software across all twenty hosts in your cluster. Your virtual machines have only ever run on six of them. What is your strongest response? A, explain that the cluster is oversized and the workload is small. B, produce the migration history showing which hosts the VMs could and did reach. C, ask for a commercial discount on the total, given the relationship. D, deploy affinity rules now and report the reduced boundary. Pause here, and ask which of these changes the size of the claim rather than the price of it.
The answer is B, produce the migration history. Evidence changes the size of the claim while a discount only changes its price, and asking for the discount quietly accepts the model that produced the number in the first place. Answer D is worth dwelling on, because it is the right thing to do and the wrong answer to this question. Affinity rules applied today shrink tomorrow's boundary and prove absolutely nothing about the period under audit. So do it, immediately, on the day. Just do not confuse it with a defence, because those are two different pieces of work with two different beneficiaries.
Guest analyst clip. I want to leave you with the smallest possible action, because this session has been full of large numbers and large numbers tend to produce paralysis rather than movement. Go and find out how long your hypervisor keeps its migration history. That is it. That is the action. It will be a number in a settings panel somewhere, it was almost certainly set to a default, and the person who set it had no idea they were setting the retention period on your audit defence. In the largest defence in our file, migration history was the single artefact that removed most of the claim. Not the contract, not the tooling, not anybody's argument. A log of which machines had run where, which existed only because somebody had never got round to reducing a default. That is an uncomfortable thing to have been lucky about. So the ask is simply to stop being lucky. Find the setting, work out what your retention actually is, and if it is ninety days, have the conversation about making it two years to match the reporting retention we covered last session. The storage cost of that is trivial. The evidentiary value of it, on the day it matters, is the difference between two very different numbers.
So, making the boundary a decision rather than an accident, five things. A named IBM footprint, meaning which clusters and which pools IBM software is permitted to live in, written down, so that permission is granted rather than inherited. A licensing line on the change form, so that host additions, cluster merges, affinity changes and activations each ask one question, does this change the licence position and who checked. Mobility history retained for two years, matched to the report retention from last session, because the two together are what an audit response is actually built from. Boundary reviewed with the quarterly close, so the cluster membership list sits next to the PVU summary, since a report is only as meaningful as the boundary it was calculated inside. And the levers priced before they are needed, so you know what dedicating a cluster or capping a pool would save, and when a project asks for headroom the answer includes a number.
Three sentences. The countable boundary is the set of hosts a workload is permitted to run on, so one eight core workload is five hundred and sixty PVUs pinned and measured, and four thousand four hundred and eighty free to move across four sixteen core hosts, with nothing about the workload itself different between those two numbers. The boundary grows through hosts added, mobility widened, clusters merged and capacity activated on demand, and not one of those is made by a licensing team, which is why the control belongs in the infrastructure change process rather than in a quarterly review. And every lever that shrinks the boundary is a configuration you already control, while every argument about a boundary in the past is settled by mobility history rather than by explanation, which is why one defence took an opening claim of eighty two million dollars to six hundred thousand on migration logs.
Homework, about an hour. List the clusters IBM software lives in, not the virtual machines, the clusters, and if that list has never been written down then writing it down is the whole exercise. Count the hosts in the largest one, multiply by the core count and the rating, and that number is your exposure on that cluster if measurement ever lapses. Find one affinity rule, asking whether any rule restricts where IBM workloads may be placed, and if not, ask what it would cost to add one in performance terms rather than in licences. Check the migration log retention, meaning how long your hypervisor keeps migration history and who set that number, because it was almost certainly not set with an audit in mind. And ask about the last expansion, finding the most recent host addition or cluster merge and asking whether anybody raised the licence question at the time, because the answer tells you whether the control exists at all.
Five guides. IBM PVU licensing explained carries the cluster worked example, the argument for why mobility settings are licensing decisions, and the case for dedicated IBM clusters. The Power Systems and AIX licensing guide covers activated cores, micropartitioning, capped shared pools, the capacity on demand trap and the posture comparison, and it is the one to read if you run Power at any scale. And the PVU sub capacity guide covers the eligible technologies, the products that cap nothing because they are full capacity only, and what a lapse costs on a partitioned host.
The eighty two million dollar audit defence case study is worth reading in full, for the cluster wide counting, what the migration history proved, and why the first sixty days went on evidence rather than negotiation. And the pharmaceutical case study gives you the factor of seven on dense clusters, and how cluster changes that were never mirrored in the tool cost forty percent of coverage. Next time, module two closes with the ILMT operating model: the five failures that create findings, and how to keep the position clean between audits. See you there.