The word covered everything from a chat window to a fully autonomous negotiator, and the useful definition turned on initiative and boundaries rather than on autonomy
A buyer cannot evaluate what they cannot define. The word arrived faster than any shared meaning did, and the confusion cost money in both directions.
Prepared by Redress Compliance · August 19, 2026 · Buyer conversations as agents entered procurement, 2024 and 2025.
Executive summary
The agents that delivered value were narrow and grounded: a price check, a terms scan, a renewal watch, each owning one job well.
The agents that frightened people were the ones described as autonomous, which is exactly the framing to avoid when you are trying to get one adopted.
Every team that understood the three rung ladder adopted faster, because they knew what they were actually buying.
Three properties make an agent safe: grounding, bounded authority and an audit trail. Miss any one and the agent is a liability rather than a hire.
What is one, and what is it not?
Software that takes a defined job, works it through multiple steps with tools and data, and returns a finished result: an analysis, a draft, a flag.
Two properties separate it from a chatbot
- It takes initiative, acting on an event or a schedule rather than on request.
- It completes a job rather than answering a question.
It is not an autonomous decision maker either
The useful definition sits in the middle: enough initiative to own a task, enough boundary that a human still approves anything committing the organization. Frame it as autonomy and you have described the thing to avoid.
What separates the three rungs?
How much of the work each one actually does, and who drives. The rungs are routinely confused, and the confusion is what makes evaluation impossible.
| Rung | Who drives | Example in procurement | What it is good for |
|---|---|---|---|
| Assistant | You ask, it answers | Look up a definition or a clause on request | Lookup, useless for coverage |
| Copilot | You drive, it assists | Prompt you during a live vendor call | Working alongside a person inside a task |
| Agent | It owns the job | Benchmark a proposal the moment it arrives, unasked | Coverage without anybody remembering to ask |
Most chat interfaces are assistants wearing a better name
An assistant answers when asked, so everything depends on somebody remembering to ask. That is the difference between a useful tool and actual coverage.
A copilot is genuinely useful and genuinely different
It works alongside a person inside a task they are driving, suggesting clauses in a review or surfacing a fact during a call. The human holds the wheel. Initiative and completion are what put the third rung above it.
The procurement platform buyer guide
What coverage a platform actually gives, where judgment still belongs to people, and how the two hand off.
Read the guide →What the buyer conversations showed
In the buyer conversations Morten Andersen had as agents entered procurement in 2024 and 2025, the word covered everything from a chat window to a fully autonomous negotiator, and the confusion had real cost. Three findings recur.
- The agents that delivered value were narrow and grounded: a price check, a terms scan, a renewal watch, each owning one job well.
- The agents that frightened people were the ones vendors described as autonomous, which is exactly the framing to avoid.
- Every team that understood the assistant to copilot to agent ladder adopted faster, because they knew what they were buying.
Teams dismissed agents as rebranded chatbots or feared them as unsupervised deal makers. Both reactions came from the same missing definition.
- Every risky clause flagged with the verbatim quote and page anchor
- Your agreements decoded into plain English before the auditor interprets them for you
- A ranked savings queue with dollar values, not license counts
What makes one safe?
Three properties, all mandatory before it goes near a real deal. These are the definition of a safe agent rather than optional refinements.
Grounding, bounded authority, audit trail
- Grounding: every output cites the benchmark, clause or invoice line it stands on. No citation, no output.
- Bounded authority: it drafts, classifies and flags. It never sends externally, signs or concedes.
- Audit trail: every action logged, every draft attributable, every decision reviewable.
The standards exist because the hazard is known
The public frameworks describing these properties, from the risk management framework to the tool connection conventions of the open protocol standards, exist precisely because an ungrounded, unbounded, unlogged agent is a documented risk rather than a theoretical one.
Bounded authority is what makes adoption possible
The teams that feared unsupervised deal makers were reassured by the boundary, not by the capability. Naming what the agent cannot do is more persuasive than describing what it can.
Where do they actually help today?
In narrow, grounded jobs rather than broad autonomous ones. The highest value and lowest risk pattern is background monitoring that needs no prompt at all.
Three patterns that work now
- Email agents: forward a proposal and get a price check or terms scan back, from the inbox the team already uses.
- Background jobs: renewal alerts at 120, 90 and 60 days, and invoice matching that runs daily.
- First reads: every new document summarized and flagged before anybody opens it.
No prompt is the point
Coverage is the thing an assistant cannot give, because it depends on somebody remembering. A background job that runs whether or not anyone thinks of it is the whole argument for the top rung. The task level breakdown sits in what agents actually automate.
Where the common framing is wrong
The common framing is autonomy: the more the agent decides on its own, the more advanced it is. We disagree.
Autonomy is the framing that blocks adoption
The agents that frightened people were exactly the ones described as autonomous, and the agents that delivered value were narrow, grounded and bounded. Vendors sell the first framing; buyers need the second.
The buyer side move is to evaluate on initiative and boundaries: what it starts without being asked, what it cites, and what it is structurally unable to do. The platform question sits in what procurement software is and the agent inventory in the procurement agents reference.
What the conversations established, 2024 and 2025
Two definitional findings rather than measured ones, and both changed what buyers could evaluate.
Taking initiative on an event or schedule rather than on request, and completing a job rather than answering a question.
Grounding with citations, bounded authority that never commits the organization, and an audit trail on every action.
Neither is a benchmark. Both are the test a buyer applies before any capability conversation is worth having.
Your first five moves
- Fix the definition before the evaluation, because a buyer cannot evaluate what they cannot define and the word covers three different things.
- Place each candidate on the ladder, since most chat interfaces are assistants and an assistant cannot give you coverage at all.
- Require grounding on every output, with the benchmark, clause or invoice line it stands on. No citation, no output.
- Bound the authority in writing, so it drafts, classifies and flags and never sends externally, signs or concedes.
- Start with background monitoring rather than negotiation. The procurement practice starts from the narrow grounded jobs, which is where the value and the lowest risk both sit.
Frequently asked questions
What is an AI procurement agent?
Software that takes a defined job, works it through multiple steps with tools and data, and returns a finished result such as an analysis, a draft or a flag.
What separates it from a chatbot?
Two properties. It takes initiative, acting on an event or schedule rather than on request, and it completes a job rather than answering a question.
Is it an autonomous decision maker?
No, and that framing is the one to avoid. The useful definition has enough initiative to own a task and enough boundary that a human approves anything committing the business.
What are the three rungs?
Assistant, which answers when asked. Copilot, which assists a person driving a task. Agent, which owns the job and starts without being asked.
Why does the ladder matter?
Because every team that understood it adopted faster. They knew what they were buying, which is impossible when one word covers all three rungs.
What makes an agent safe?
Grounding so every output cites its source, bounded authority so it never sends or signs, and an audit trail so every action is logged and reviewable.
Are those optional?
No. They are the definition of a safe agent rather than refinements, which is why the public frameworks describe them explicitly.
Where do they help today?
In narrow grounded jobs: email price checks and terms scans, background renewal alerts and invoice matching, and first reads of every new document.
Why is background monitoring the best pattern?
Because coverage is the thing an assistant cannot give. A job that runs whether or not anybody thinks of it is the whole argument for the top rung.
How should a buyer evaluate one?
On initiative and boundaries: what it starts without being asked, what it cites, and what it is structurally unable to do.