The assistant is the one number you have, and it cannot predict the one you are buying
Joule agents and the Joule assistant draw on the same AI Unit balance, so buyers reasonably use the assistant to estimate what agents will cost. That estimate fails, and not only because an agent does five to ten times the work. It fails because a human stops asking and a schedule does not.
Prepared by Redress Compliance · August 11, 2026 · SAP advisory. Based on SAP Business AI engagements benchmarked in 2024 and 2025.
Executive summary
A Joule agent consumes roughly 5 to 10 times an interactive prompt per task, and both draw on the same AI Unit balance. The shared balance is what makes the assistant look like a reasonable basis for a forecast.
It is not, because the two differ in more than magnitude, and a buyer who scales the assistant number by ten still lands well short of the production figure.
Interactive Joule sits mostly inside bundled Base while agents sit in paid Premium. That is why the assistant rarely appears as a meaningful bill and why buyers conclude Business AI is inexpensive.
The conclusion is correct for chat and wrong for agents, and the estate has no way to discover this from its own usage data before it commits.
Pilot cost is a poor predictor of scheduled production consumption.
A pilot agent run a few times a day is a fundamentally different object from a scheduled fleet running continuously, so the pilot measures the multiplier per task and tells you nothing about the two multipliers that follow it: run frequency and fleet size.
The controls that work are central schedule approval and a contractual cap. Both are governance rather than engineering, because the consumption is driven by how many agents are scheduled and how often, and those are decisions made in project meetings rather than settings tuned by an administrator.
How each one consumes
| Dimension | Joule assistant | Joule agent |
|---|---|---|
| Work per task | One turn | 5 to 10 actions |
| Trigger | Human prompt | Schedule or event |
| Tier | Mostly bundled Base | Premium, draws AI Units |
| Ceiling on demand | Human attention | None once scheduled |
| Budget risk | Low | High, compounding |
The row that matters is not the first one, it is the second. Work per task is the difference everyone models: five to ten actions instead of one turn, so multiply by ten and move on. The trigger row is the difference nobody models, and it is the one without an upper bound.
A human prompt is self limiting because people only ask so many questions in a working day, so assistant consumption tracks headcount and attention. A schedule has no such limit, and neither does the number of agents an organisation decides to schedule.
The metering mechanics sit in the AI Units explainer.
Controlling run rate before it reaches production
- Route every agent schedule through a central approval, because run frequency is decided in project meetings rather than set by an administrator, and it is the multiplier with no natural ceiling.
- Model the fleet, not the agent, since consumption is actions per task multiplied by frequency multiplied by the number of agents, and buyers routinely model only the first term.
- Treat the pilot as a measurement of the multiplier only, as a pilot run a few times a day cannot tell you what a continuously scheduled fleet will draw.
- Negotiate a contractual cap on AI Unit consumption, because a shared balance drawn by unattended processes is the case where a commercial limit does work a technical control cannot.
- Separate the Base and Premium draw in your forecast, so that bundled interactive usage stops flattering the number you are using to size the paid tier.
The SAP AI Units renewal playbook
The buyer side moves on SAP Business AI, AI Units metering, Joule agent consumption, and the use based renewal default.
Get the playbook →A change of trigger, not a change of size
The five to ten times figure is accurate and it is also the least dangerous part of this. A known multiplier is something a finance team can handle: measure the assistant, multiply, add margin.
What breaks the forecast is that the multiplier is not the only thing that changes when an estate moves from chat to agents, and the other change has no arithmetic attached to it. The assistant is triggered by a person, which means its consumption is bounded by human attention.
There are only so many questions a workforce asks in a day, so assistant usage scales with headcount and settles into a stable band. An agent is triggered by a schedule or an event, and a schedule has no equivalent ceiling.
Nobody gets tired, nothing throttles it, and the run frequency is set by whoever configured the job rather than by demand.
Then the same thing happens at the fleet level: the number of agents an organisation runs is a product decision, not a demand signal, and it tends to grow as teams find new candidates for automation.
So the true consumption is actions per task multiplied by run frequency multiplied by fleet size, and only the first of those three is what buyers measured in the pilot. This also explains why the assistant feels free and why that feeling is so misleading.
Interactive Joule mostly sits inside bundled Base, so the estate genuinely does not see a meaningful bill from it, and it reasonably concludes that Business AI is inexpensive.
That conclusion is correct for chat and wrong for agents, and critically, the estate cannot discover the error from its own data before committing, because the only usage it has is the usage that does not predict.
The practical consequence is that the controls have to be governance rather than tuning.
Central approval of agent schedules works because it puts a person in front of the multiplier that has no ceiling, and a contractual cap works because a shared balance drawn by unattended processes is exactly the situation where a commercial limit does what a technical setting cannot.
The full licensing model sits in the Joule and AI Units pillar.
- Percentile standing for your exact deal size and industry, from real closed transactions
- Scenario simulation before the call: test alternative terms and see the financial impact of each
- A negotiation playbook, talking points, and a two page executive brief on day one
What we saw across Joule agent and assistant engagements, 2024 and 2025
Across the SAP Business AI engagements benchmarked in 2024 and 2025, the largest budgeting error was treating a Joule agent like a bigger Joule prompt:
What an agent draws against a single interactive prompt, because it reads context, reasons, calls systems, and writes results.
Buyers modelled actions per task and left run frequency and fleet size out, which is where production consumption actually lives.
Three patterns recurred: the assistant felt free because it sat inside bundled Base, the first pilot looked cheap because it ran a few times a day, and the scheduled production fleet drew the balance five to ten times faster than modelled.
The buyer side move is central schedule approval plus a contractual cap. The wider library sits in the SAP practice.
Your first five moves
- Stop forecasting agents from assistant usage, because interactive Joule sits mostly in bundled Base and is the one number you have that cannot predict the one you are buying.
- Model all three multipliers, actions per task, run frequency, and fleet size, since the pilot measures only the first and the other two are the ones without a ceiling.
- Put central approval in front of every agent schedule, as run frequency is decided in project meetings and is the term that grows without a demand signal to check it.
- Negotiate a contractual cap on AI Unit consumption, because unattended processes drawing a shared balance is the case where a commercial limit beats a technical control.
- Re baseline after the first scheduled fleet runs, then renegotiate on real consumption. The SAP practice runs the model and the renewal together.
Frequently asked questions
How much more does a Joule agent consume?
Roughly 5 to 10 times an interactive prompt per task, because an agent completes multi step work: it reads context, reasons, calls systems, and writes results, and each step draws from the balance. Both the assistant and agents draw on the same AI Unit balance.
Why does the assistant set a misleading baseline?
Because interactive Joule mostly sits inside the bundled Business AI Base tier, so it rarely produces a meaningful bill. Buyers conclude Business AI is cheap, which is correct for chat and wrong for agents, and the estate cannot discover the error from its own usage data before committing.
Is the five to ten times multiplier the main risk?
No, it is the most manageable part. A known multiplier can be modelled. The larger risk is the change of trigger: an assistant is bounded by human attention while a scheduled agent has no such ceiling, and neither does the number of agents an organisation decides to run.
Why does a pilot understate production cost?
A pilot agent run a few times a day measures the per task multiplier and nothing else. Production consumption is that multiplier times run frequency times fleet size, and a scheduled fleet running continuously is a different object from a pilot rather than a larger one.
Which tier do agents draw from?
Agents sit in paid Premium and draw AI Units, while interactive Joule largely sits in bundled Base. Both draw against the same balance, which is what makes the bundled interactive usage flatter the forecast used to size the paid tier.
What controls actually work?
Central approval of agent schedules and a contractual cap on consumption. Both are governance rather than engineering, because run rate is set by how many agents are scheduled and how often, and those are decisions taken in project meetings rather than settings an administrator tunes.
How should the forecast be built?
As three multipliers rather than one: actions per task, run frequency, and fleet size. Separate the Base and Premium draw so bundled interactive usage does not flatter the number, and re baseline once the first scheduled fleet has run in production.