HomeTraining AcademyIBM Licensing MasterySession 8
IBM Licensing Mastery · Module 2 – PVU, sub capacity and ILMT · Session 8 of 20 · 23:49

ILMT

Deployment, the scan cadence, the quarterly report, the two year retention, and the five failures that create findings. Three knowledge checks along the way, and 4 clips from a senior licensing analyst.

What you will be able to do after this session

  • 1Separate installed from operated. In 18 of 25 remediation projects the tool was installed and technically working, and what failed was operations after go live.
  • 2Hold the cadence stack. Capacity scans every 30 minutes, software scans weekly and never past the 30 day floor, hypervisor collection live, reports quarterly, retention two years.
  • 3Find the connection that severs silently. Without hypervisor visibility the sub capacity calculation cannot be made, and the credential that provides it breaks on rotation without alerting anybody.
  • 4Audit your own catalog. Stale catalogs inflated PVU counts by 10 to 25 percent, which is money you are paying for a classification error rather than for software.
  • 5Run a quarterly close. Generate, review, sign, export, archive, on a named date. Audits read the report trail rather than the install date.

How the session works

This is a taught session, not a talking head. The instructor works through analyst grade slides, and three times the video stops on a question with four options on screen. Pause, commit to an answer, and the next slide explains which option is right and why each of the others is wrong. 4 times in the session the frame splits and a senior licensing analyst gives the view from inside real IBM negotiations, and the instructor picks the clip apart when the slides return.

Homework before session 9, about one hour

  • 1Check the hypervisor connection. Is it live, and when was it last verified? If nobody can answer the second question, the answer to the first is probably not what you hope.
  • 2Read the poll cadence. Find the configured capacity scan interval. If it is over 30 minutes, find out who changed it, when, and whether the change is documented.
  • 3Age your software scans. Find the oldest software scan age across eligible hosts. Anything past 21 days is a host heading for a gap.
  • 4Count the unclassified. How many discovered components are sitting unclassified, and how old is the oldest? That queue is where the 10 to 25 percent lives.
  • 5And time an evidence request. Ask for the PVU summary and entitlement reconciliation for the quarter before last, and time how long it takes to arrive. That is your audit response time.

Session transcript

The full narration of this session, section by section, for reading and reference. Guest analyst clips are marked.

Welcome and objectives 0:02

Welcome to session eight. Last time we established that sub capacity rests on four obligations, and that three of the four are carried by a single piece of software. Today is that software. And I want to start by narrowing what this session is about, because it is not a tutorial. You will not learn to install anything here. What you will learn is the difference between a tool that is installed and a tool that is operated, because that difference is where essentially all of the money is. In eighteen of twenty five remediation projects the tool was installed and technically working. What failed was operations after go live. Three knowledge checks. Let's begin.

Five objectives. First, separate installed from operated, for the reason I just gave. Second, hold the cadence stack, which is capacity scans every thirty minutes, software scans weekly and never past the thirty day floor, hypervisor collection live, reports quarterly, and retention for two years. Third, find the connection that severs silently, because without hypervisor visibility the sub capacity calculation cannot be made at all, and the credential that provides it breaks on rotation without alerting anybody. Fourth, audit your own catalog, because stale catalogs inflated PVU counts by ten to twenty five percent, and that is money you are paying for a classification error rather than for software. And fifth, run a quarterly close, meaning generate, review, sign, export and archive on a named date, because audits read the report trail rather than the install date.

Installed is not operated 1:53

Four numbers. Eighteen of twenty five, the remediation projects where the tool was installed and technically working and operations after go live were what failed. Ten to thirty percent, the share of the estate that coverage drift had left outside sub capacity protection, unnoticed. Ten to twenty five percent, how much a stale catalog inflated the PVU count, which I would underline because that is money paid for a classification error rather than for software. And two quarters, which is how long the stack takes to decay when nobody owns it, because agent drift on new hosts went uncaught for one to two quarters on most estates. The note underneath names the cause behind all four. The pattern behind the pattern is ownership. The tool belonged to nobody: installed by a project, inherited by no team, and discovered broken by an audit.

Guest analyst clip. The sentence about the tool belonging to nobody is the truest thing in this session, and I want to describe how it happens, because nobody decides it. There is a project. It has a budget, a manager, a deadline and a success criterion, which is usually that the tool is installed and producing a report. The project achieves that, entirely competently, and then it does what projects do, which is end. And at the moment it ends, the tool passes to nobody, because there was never a line in anyone's job description that said this is yours now. The infrastructure team think it is a licensing tool. The licensing team think it is infrastructure software. Both are being reasonable. And it keeps running, because that is what software does, and it keeps producing reports, which is what makes this so dangerous, because the visible output continues long after the substance has drifted. So when I am asked what the single highest value intervention is in an IBM estate, my answer is not technical at all. It is to put one name against this thing, and to give that person a financial mandate rather than an operational one, because what they are protecting is a number, not a service.

One name, with a financial mandate rather than an operational one, because what they are protecting is a number and not a service. Now, what it is they would be protecting.

What you are deploying 4:12

Four parts, and where each one goes wrong. The agent, which sits on every host running eligible products and scans it, and which goes missing on new hosts and gets broken quietly by operating system patching. The server, the console and report engine and scheduler sized for the estate, whose version ages, and an aged version hands IBM an argument that the tooling was invalid. The database, which holds the scan and report history, two years of it, and where disk pressure purges history and upgrades lose it altogether. And the hypervisor link, a read only credential that supplies physical capacity, which severs silently on credential rotation or a vCenter migration. The note is the one I would take away. The retention policy on that database is the retention policy on your audit defence. Treat it like an audit log rather than like application data, because those two things get very different treatment when somebody needs disk space.

Knowledge check 1 5:17

Knowledge check one. Your ILMT server is current, agents are on every host, and reports run on schedule. The vCenter credential was rotated eight months ago and nobody reconnected it. What is the risk? A, none, because the agents are reporting and the reports are generating. B, the sub capacity calculation has no physical capacity to compare against, so the eligibility itself is in question. C, only that the reports will be a little less accurate. D, the agents will stop scanning within thirty days. Pause here, and ask what a sub capacity number actually compares.

The answer is B. Sub capacity maths compares virtual cores to physical capacity, so without hypervisor visibility the calculation fails, and your eligibility goes with it. Answer A is what the monitoring dashboard will tell you, and there is a general lesson in that. Every component with an alert attached is healthy, and the one component nobody attached an alert to is the one that broke. That is not a coincidence, it is the definition of the problem. This is the single most missed step in ILMT deployments, and the credential is also the longest lead time item to obtain in the first place, which is why it is worth protecting once you have it.

The cadence stack 6:49

Five clocks, all running at once. Capacity scans every thirty minutes, unbroken, because a polling cadence longer than thirty minutes breaks the sub capacity entitlement, and custom settings made for good performance reasons are the usual cause. Software scans weekly by default and never past the thirty day contractual minimum, with an alert when any eligible host's software scan age passes twenty one days, which leaves you a week of margin. Hypervisor collection every twelve hours in central mode, so the physical capacity picture stays current as the estate moves underneath it. Reports quarterly at minimum, which is the contractual floor rather than the target, and monthly costs almost nothing more while giving you three chances a quarter to notice a problem. And retention for two years, twenty four months minimum, monitored for disk and database growth, because history purged under disk pressure is history you cannot get back.

The connection that severs silently 7:55

So let me take the hypervisor connection properly, because it deserves its own slide. What it does is supply the physical capacity of each host, which is the denominator in every sub capacity calculation the tool performs. Why it is missed at build time is that it needs a read only hypervisor credential from a different team, and those requests are the longest lead time item in the whole deployment, so they get deferred and then forgotten. How it breaks afterwards is credential rotation and vCenter migrations, and it breaks silently, because nothing in the tool fails loudly when a connection simply stops returning data. What it looks like when broken is agents reporting, reports generating, dashboards green, and a sub capacity number that no longer has anything to be sub of. And the control is to re verify after any vCenter migration or credential rotation, and to put that verification into the change process rather than into somebody's memory.

Guest analyst clip. I have developed a rule of thumb over the years that I think generalises well beyond this tool, so let me offer it. In any system, the components most likely to fail undetected are the ones that fail by going quiet rather than by going wrong. A crash generates a log entry. A timeout generates an error. But a connection that simply stops returning data generates nothing at all, and every downstream process carries on happily with whatever it last knew. That is the hypervisor link exactly. Nobody built an alert for it, not through carelessness, but because there is no event to attach an alert to. The absence of data is not an event. So the discipline has to be a positive check rather than a negative one. Not, has anything broken, but, when did we last confirm this is working, and can somebody show me the date. And I would put that question into your change process at three specific moments: after any hypervisor migration, after any credential rotation, and after any upgrade of the tool itself. Three moments, one question each. It takes minutes and it closes the gap that produces the most surprising audit findings I see, the ones where everything looked healthy for two years.

The absence of data is not an event, so the check has to be positive rather than negative. Not has anything broken, but when did we last confirm this works, and who can show me the date.

Knowledge check 2 10:28

Knowledge check two. Your quarterly PVU report shows a number twenty percent higher than your own estimate. What is the most likely cause? A, IBM has changed the value unit table. B, the catalog and bundle map are stale, so components are counted standalone rather than inside their bundle. C, your estimate is simply wrong and the tool is authoritative. D, the agents are double counting hosts. Pause here, and ask what the tool has to be told rather than being able to work out for itself.

The answer is B. Stale catalogs inflated PVU counts by ten to twenty five percent in remediation work, and a bundle component left untagged is treated as a standalone product. Answer C is the assumption that costs the most money, and it is worth dwelling on, because the tool feels authoritative in exactly the way a Passport Advantage export felt authoritative back in session five. It is precise, it is generated by software, and it is entirely dependent on how somebody configured it. Configuration must mirror your real deployments, and this is the one place in the whole module where the number moves in your favour when you do the work.

The catalog and the bundle map 11:54

So, the catalog, which is where the hidden money sits. The tool reports what it was told, because bundling, component exclusions and product assignments all change the PVU result and none of them are things a scan can work out by itself. Unclassified discoveries rot into findings, so classify new discoveries within thirty days, because an unclassified component sitting in the tool is an untagged product sitting in an audit. The bundle map is the expensive one, since a component entitled inside a bundle but left untagged is counted as a separate product, which is the same double count from session five wearing different clothes. Review it quarterly, in the same close as the report, because deployments change faster than anybody updates a classification. And notice the direction of travel here. Almost everything else in this module protects you from a larger bill. Catalog work reduces the bill you are already paying.

Guest analyst clip. I want to make an argument for why the catalog work gets done, because in my experience knowing about it is not enough. Everything else in this area is insurance. You do it so that a bad thing does not happen, and insurance is famously difficult to get funded, because the return is invisible when it works. The catalog is different, and I think it is worth selling internally on exactly that difference. A stale catalog inflates the count by ten to twenty five percent, and you are paying for that inflation right now, this quarter, in real support renewals on entitlements you did not need to hold. Fix it and the saving appears in a number somebody already tracks. So my advice to anybody trying to get this funded is: lead with the catalog. Do not open by explaining audit risk, because audit risk is a story about the future and the future has many competing stories. Open by saying we are currently overstating our own consumption by something in the region of a fifth, and here is what that is costing us annually. Then, once you have the attention and the owner and the quarterly rhythm, all of the insurance work comes along with it, because it is the same person doing the same close. The unglamorous protection rides in on the back of the visible saving.

Lead with the catalog, because the saving is visible and this quarter. The protection then rides in on the back of it, done by the same person in the same close.

The five failures 14:22

The five findings, and the control that prevents each one. The agent missing on a host, where a new host stood up without the agent breaks sub capacity for the products on it, and the control is the agent in the gold image so hosts arrive covered rather than getting covered. Polling cadence over thirty minutes, where a custom poll set to an hour drops the entitlement, and the control is the cadence pinned, the configuration locked and a change history you can show. Retention under twenty four months, which is disk pressure, history purged, audit window missing, and the control is a growth alarm on the database rather than a cleanup script somebody wrote once. The bundle map incomplete, where a component left untagged is audited as standalone, and the control is the quarterly review, which takes an hour and moves the number in your favour. And the quarterly review missing, which is no summary, no sign off, no archive, and an audit that revokes sub capacity status, and the control is a date in the calendar with a name against it.

The quarterly close 15:36

So what does a quarterly close actually contain. The PVU summary, meaning peak PVU per product per host across the period, and this is the critical exhibit and the one an auditor reads first. The entitlement reconciliation, consumption set against what you own, which is what turns a technical report into a licence position and is the second critical item. The server and partition list and the bundle map, which is what was scanned and how it was classified, and both matter because that is how a reader checks whether the summary can be trusted at all. The anomaly log, covering scan failures, coverage gaps and their resolutions, and investigate every failure within ten business days because the documentation is itself audit evidence. And the signature, software asset manager, infrastructure lead, CIO designate, then export the snapshot and archive it the same day, and the quarter is closed.

Knowledge check 3 16:42

Knowledge check three. An audit letter arrives. Your ILMT has been current and complete for the last three quarters, and the four before that are missing. Where do you stand? A, clean, because the tool is current now. B, exposed for the four missing quarters, because the audit tests the report history rather than the current state. C, exposed for everything, because the trail is broken. D, clean, because three quarters establishes a pattern of compliance. Pause here, and ask which periods you can evidence and which you cannot.

The answer is B, exposed for the four missing quarters. Audits test two years of quarterly report history rather than the install date, and the exposure is bounded to the periods you cannot evidence. Answer C is worth naming carefully, because it is very often how the opening claim is framed, and it is not correct. Three good quarters are three quarters you do not pay for. So the first move on receiving a letter is not to negotiate and it is not to panic, it is to assemble every quarter you can actually evidence, because each one you produce is a period that comes off the claim.

Guest analyst clip. What happens in the first two weeks after an audit letter decides more than anything that happens in the negotiation, and it is almost always the same mistake. The letter arrives, it contains a number, and the organisation's instinct is to react to the number. Is it fair, is it survivable, who do we know at IBM, should we get lawyers. All understandable and all premature, because the number is not a fact, it is a model, and the model was built on assumptions about periods the auditor could not see. Which means the highest value work available to you is not argumentative at all, it is archaeological. Go and find every quarter you can evidence. Every report in a folder, every export somebody saved, every archived database backup that still has the history in it. Each one you produce removes a period from the claim, and removing periods shrinks the model rather than disputing it. And there is a real difference between those two things, because negotiating a discount accepts the model and asks for mercy, while producing evidence replaces the model with a smaller one. I have watched extremely large opening claims collapse this way, without anybody raising their voice, purely because somebody spent two weeks looking in the right places before anybody spent an hour arguing.

The operating model 19:19

So, running it like a control rather than a project, five things. One named owner, because without one the stack decays within two quarters, and that owner needs a financial mandate for the reason we opened with. Coverage held at ninety eight percent or better, with every gap treated as an incident rather than a backlog item, and new hosts enrolled inside seven days of provisioning. The quarterly close in the calendar, dated before the quarter starts, with the signatures named, and the archive step treated as the part that actually matters. Version and credential checks in change control, so hypervisor migrations, credential rotations and tool upgrades each raise a verification step, because all three break things silently. And an evidence pack that can be produced in days, two rolling years and findable, so that an audit letter starts with assembling what you have rather than discovering what you do not.

Recap 20:22

Three sentences. In eighteen of twenty five remediation projects the tool was installed and technically working, and what failed was operations after go live, which is why installed is a ticket and operated is the defence. The cadence stack is capacity scans every thirty minutes, software scans weekly and never past the thirty day floor, hypervisor collection every twelve hours, reports at least quarterly and retention for two years, and the hypervisor connection is the one that breaks silently, because nothing alerts when a credential simply stops returning data. And catalog and bundle map work is the only part of this module that reduces a bill you are already paying, since stale catalogs inflated PVU counts by ten to twenty five percent, while everything else here protects you from a larger one.

Homework 21:18

Homework, about an hour, and this one is a set of five questions with uncomfortable answers. Check the hypervisor connection: is it live, and when was it last verified, because if nobody can answer the second question the answer to the first is probably not what you hope. Read the poll cadence, finding the configured capacity scan interval, and if it is over thirty minutes find out who changed it, when, and whether the change is documented. Age your software scans, finding the oldest software scan age across eligible hosts, because anything past twenty one days is a host heading for a gap. Count the unclassified, meaning how many discovered components are sitting unclassified and how old the oldest one is, because that queue is where the ten to twenty five percent lives. And time an evidence request, by asking for the PVU summary and entitlement reconciliation for the quarter before last and timing how long it takes to arrive. That number is your audit response time, and you have just measured it for free.

Further reading 22:29

Five guides. The ILMT deployment guide covers the six step deployment, the cadence stack, the hypervisor credential lead time and the approved alternatives, and it is the one to hand an infrastructure team. The deploy and configure guide gives you the three tier architecture, the report contents ranked by audit weight, the five common findings and the disaster recovery trap. And the comprehensive pillar treats the tool as a regulated control, covers the retention question properly, and makes the case that most failed audits are deployment failures rather than licensing decisions.

The sub capacity readiness guide is a check you can run against your own estate with the full capacity comparison worked through. And the pharmaceutical case study is worth reading for scale, a fourteen week remediation across roughly three thousand hosts, and what forty percent coverage loss turned out to be worth once somebody looked. Next time, virtualization and capping: eligible technologies, partitions and shared processor pools, VMware clusters, and how one host becomes the whole estate. See you there.

Learning the playbook and want it applied to your numbers? We work on contingency: 25% of what we save you. Nothing saved, nothing paid.
Review my deal