Contents
Key takeawaysExtraction versus summaryHow the pipeline worksWhat our projects showedWhere extraction failsTesting accuracyA trustworthy designTalking to tool vendorsWhat to do nextFAQA summary tells you roughly what an agreement says. Extraction tells you exactly what it says, in fields you can search, match and monitor, as long as a person confirms the fields that carry money and dates.
- Extraction produces data. Each field, such as a renewal date or an uplift cap, is tied to its source clause, so it can drive calendars and invoice checks.
- Three failure points. Errors cluster in amendments, order forms that override the master, and terms defined by reference, which is where the operative price usually lives.
- Confirmation is the stage that matters. Projects that skipped human confirmation put a wrong renewal date into a live calendar.
- Page references make review fast. With a link to the source clause, checking a field takes seconds instead of a reread.
- Test on your own contracts. A headline accuracy figure from a demonstration set averages away the errors in amended contracts.
- Index only confirmed fields. Unconfirmed values in search results and calendars get trusted by people who never saw the confidence signal.
AI contract extraction reads a signed agreement and turns it into a record: renewal date, notice window, uplift cap and liability cap, each with its source page. On clean master agreements the current models do this well. The trouble starts in the documents that change the master, which are the ones that set what you pay.
This guide covers how extraction works, where it breaks, how to test a tool on your own contracts, and how to design the review step so a wrong date never reaches your renewal calendar.
What is AI contract extraction, and how is it different from a summary?
AI contract extraction converts an unstructured document into structured records: named fields with values, and clause positions with locations. The output is a set of fields, each tied to the clause it came from, which other systems can read. A summary is written for a person to read once.
Why can you not run anything off a summary?
Summarization paraphrases, and a paraphrase cannot be queried. Extraction produces a renewal date field, an uplift cap field and a liability cap field, each pointing back to its source clause. Ask a folder of summaries which contracts renew next quarter and someone has to reread every one of them.
You can run a renewal calendar off the renewal date field and an invoice check off the uplift cap and the rate card. That is the reason to care about the distinction before you trust any output: every downstream use assumes the field is right.
Which fields should a software contract record hold?
Start with the fields that carry money and deadlines, then add the legal positions. Each field below also has a typical way of going wrong, which tells your reviewers where to look.
| Field | What it drives | Where extraction tends to slip |
|---|---|---|
| Renewal date and term | The renewal calendar | Reads the original term and misses an amendment that extended it |
| Notice window and auto renewal | The last day you can exit or renegotiate | Notice counted from the wrong date, or auto renewal set in an order form |
| Rate card and unit prices | Invoice matching | Master price captured while an order form set a different one |
| Uplift cap | Renewal quote checks | Cap defined in an amendment or limited to named products |
| Termination rights | Exit options during the term | Convenience rights granted in one order but not the master |
| Liability cap | Risk reviews and legal sign off | Cap expressed as a multiple of fees defined elsewhere |
| Governing documents | Which terms actually apply | Online terms incorporated by reference are never read |
How does an AI contract extraction pipeline work?
A working pipeline has four stages: read, propose, confirm and index. Buyers usually underestimate the third, because it is the only stage that is not automated and the only one that turns plausible output into data you can rely on.
| Stage | What happens | Why it matters | Where it fails |
|---|---|---|---|
| Read | Parsing turns the document into machine text | Scan quality decides everything downstream | Poor scans, unusual layouts |
| Propose | The model proposes fields and positions with confidence | Where the intelligence lives | Amendments and overriding order forms |
| Confirm | A human reviews proposals against the cited source text | Turns plausible output into trustworthy data | Skipped entirely |
| Index | Clause level indexing makes every contract searchable | Enables one question across every contract | Indexing unconfirmed fields |
Why does every field need a page reference?
Every extracted field should link to the exact page and clause that supports it. With that link, the reviewer opens the cited clause and compares it with the proposed value. Without it, the reviewer searches the whole contract for the date, and on a large backlog the review stalls before it is finished.
The reference also outlives the project. When a negotiator quotes the uplift cap to a vendor, or an auditor asks where a term came from, the answer is one click away instead of buried in a shared drive.
How should confidence route the review?
Good extraction attaches a confidence signal to each field, so high value contracts and low confidence fields reach a human first. Language models do not return a calibrated score per field on their own. Tools derive one, for example by checking that the cited text contains the value or that a notice date falls before the renewal date.
The model layer documented by makers such as one provider and another is capable, and capable is not the same as verified. The providers' own features show where that line sits.
What do citations and output schemas guarantee?
Anthropic's API can return citations with page numbers for PDF documents, which gives a tool a ready source reference. OpenAI's Structured Outputs makes the response follow a JSON schema you supply. A schema guarantees a renewal date field exists and holds a date; it says nothing about whether the date is right.
What did our 2024 and 2025 extraction projects show?
In the extraction projects I advised on in 2024 and 2025, accuracy on clean masters was excellent and the errors clustered where the money is. The same findings came up project after project.
- Page references set the speed of review. With them, a confirmation took seconds. Without them, it meant rereading the contract.
- Routing by confidence and value paid off. Sending high value and low confidence items to reviewers first put scarce human time where it changed outcomes.
- Testing on real contracts built trust. Teams that measured accuracy on their own messy contracts trusted the output. Teams that measured on the demo did not, and they were right not to.
- Skipping confirmation broke the calendar. Projects that trusted extraction with no confirmation step shipped a wrong renewal date into a live calendar. The cause was the missing stage, since the model's output reached the calendar unchecked.
These are design findings, not a vendor comparison. They describe the test to apply before you trust any calendar or invoice check built on extracted data.
Where does AI contract extraction fail?
It fails in three predictable places: amendments, order forms that override the master, and terms defined by reference. Those are exactly where the commercial terms live, which is why a small error rate overall can still produce a large error in money.
Amendments that modify an old master
An amendment changing a master signed years earlier is where extraction is weakest and where a repriced term often sits. The model has to reconcile two documents written under different assumptions. Amendments are often written as instructions, such as "the pricing schedule is deleted and replaced with the following", which mean nothing until applied to the master.
Read the master alone and you get the superseded value, and read the amendment alone and you get a change with no base. The record needs the value as amended, plus a note of which document supplies the current value. Chains of several amendments, and amended and restated agreements that replace everything before them, need the same treatment.
Order forms that override the master
An order form overriding the master for one purchase can carry the operative price, and extraction may attribute it to the wrong document entirely. A discount, a different renewal term or a waived uplift cap can live in the order form and nowhere else. The negotiations that produce these documents are covered in the AI provider negotiation series.
Which value applies depends on the order of precedence clause. Many software masters say the order form prevails for that order, while others let the master win unless the order form expressly overrides it. Link each order form to its master, extract its fields separately, and let the record show which terms apply to which products.
Terms defined by reference
When the real term lives in a linked policy or an appendix, extraction can miss it completely. Nothing in the document being read says the number is somewhere else. Large vendors do this by design: Microsoft's Product Terms form part of the customer's agreement, sit online and are generally updated on the first of each month.
The AWS Customer Agreement incorporates the Service Terms and the Acceptable Use Policy by reference, and AWS can modify them by posting a new version. For these contracts, extract the reference itself as a field (document name, location, version or date) and file a dated copy of the version that applied when you signed.
| Failure point | What goes wrong | What the record should hold | Reviewer check |
|---|---|---|---|
| Amendment to an older master | Superseded price or term reported as current | Value as amended, with the amending document named | Open every amendment linked to the master |
| Order form overriding the master | Price or renewal term attributed to the wrong document | Order form fields held separately, by product | Read the order of precedence clause |
| Term defined by reference | Field left blank or filled from the wrong text | Referenced document, location and version date | Confirm a dated copy is on file |
How do you test AI contract extraction accuracy on your own contracts?
Build a test set from your own messiest contracts and score every field by document type. A single accuracy figure across a clean demonstration set hides the errors that cost money, because it averages the easy documents with the hard ones.
Which contracts belong in the test set?
- Your oldest master with the longest chain of amendments.
- Order forms that set a price, term or cap different from the master.
- Contracts that incorporate online terms by reference.
- Poor scans and documents with unusual layouts, since the read stage fails there first.
- A handful of clean masters, as a control.
How does a headline figure hide the problem?
Take a hypothetical test: 40 contracts, 12 fields each, 480 fields in all. The tool gets 24 fields wrong, a headline accuracy of 95 percent. Now split the result by document type.
| Group | Contracts | Fields | Wrong | Error rate |
|---|---|---|---|---|
| Clean masters and simple orders | 30 | 360 | 4 | About 1.1 percent |
| Masters with amendments or overriding order forms | 10 | 120 | 20 | About 16.7 percent |
| All contracts | 40 | 480 | 24 | 5 percent |
The 95 percent sounds safe. One field in six is wrong in exactly the contracts that carry repriced terms. Score by document type and by field, and weight review toward the group where the errors sit.
What does a trustworthy extraction design look like?
A trustworthy design uses confirmation weighted by risk and value, with page references so review is fast. It assumes accuracy will be imperfect and adds a stage that catches the errors before they reach a calendar, an invoice check or a negotiation.
Three design rules
- Reference every field to its page and clause, and treat any field the tool cannot cite as unconfirmed.
- Route by confidence and value. Every date and amount in a contract with a renewal coming up goes to a reviewer, whatever its score.
- Never index unconfirmed fields. Search, the renewal calendar and invoice checks should only see values a reviewer has signed off.
Why we would not pick an extraction tool on headline accuracy
The common advice is to benchmark extraction vendors on headline accuracy and pick the highest number. We disagree. Accuracy on clean masters was excellent everywhere, so the figure separates little. The errors clustered on amendments, overriding order forms and terms defined by reference, and a headline figure measured on clean documents tells you nothing about any of those.
Choose on results from your own messiest contracts instead, scored by document type, and on how well the tool supports confirmation. The uses that depend on this sit in contract management, invoice reconciliation and deal diligence.
A searchable wrong answer is worse than no answer, because someone downstream will trust it without ever seeing the confidence signal.
Governance points the same way
The governance expectations are set out in the risk management framework and the European regulatory framework. Both put weight on documentation and human oversight of AI output.
The EU AI Act makes human oversight a legal duty for systems it classes as high risk, and a contract extraction tool is unlikely to fall in that class. A pipeline with page references and a recorded confirmation step still produces the evidence either text asks for.
What will an extraction tool vendor tell you, and how should you answer?
Expect the sales pitch to lead with accuracy and speed. Each line below has a reply that brings the conversation back to your contracts and your review process.
- "Our accuracy is the best in the market." Ask for field level results on your own test set, split by document type, with the amended contracts reported separately.
- "The model handles amendments automatically." Pick your worst amendment chain and ask the tool to show the current value of each field and which document it came from.
- "Human review is optional." Keep it on for every date and every amount. Ask how the tool records who confirmed each field and when.
- "Load everything and it becomes searchable." Ask whether unconfirmed fields can be held out of the index and the renewal calendar until a reviewer signs them off.
What should the contract with the tool vendor say?
Treat the purchase like any other software deal, because it will come up for renewal too. These are the terms we would ask for.
- Data export. Export of every field with its page reference in a standard format, so the data leaves with you.
- Pricing unit. If the tool prices per document or per page, a definition of whether each amendment and order form counts separately, since the contracts that need most review also produce the most documents.
- Training exclusion. A written commitment that your documents are not used to train models.
- Deletion on exit. Deletion of documents and extracted data within a stated period after termination, confirmed in writing.
- Confirmation audit log. A record of who confirmed each field and when, included in the export.
The data terms are covered in our guide to AI procurement data security, and the demo questions in what to ask in an AI procurement demo.
What to do next
- Build the test set first. Use the checklist above and score results by document type before you look at any vendor's headline figure.
- Require a page reference on every field. Make it a condition of purchase, and check it on the amended contracts in your test set.
- Insist on a confirmation stage weighted by value and confidence. Make it a required step in the tool's workflow so it survives deadline pressure.
- Watch the three failure points by name. Link amendments and order forms to their masters, and capture dated copies of terms defined by reference.
- Keep unconfirmed fields out of the index. Only confirmed data should feed the renewal calendar, invoice checks and search.
- Get help where the money is. Our advisory practice runs the confirmation stage as part of the pipeline, focused on the contracts with a renewal or dispute coming up.
Frequently asked questions
What is contract data extraction, and how is it different from a summary?
Contract data extraction converts an unstructured document into structured records: named fields with values and clause positions with locations, tied back to their source. A summary paraphrases and cannot be queried, while an extracted renewal date or uplift cap can run a calendar or an invoice check directly.
Where does AI contract extraction fail, and why are those failures dangerous?
It fails on amendments that modify an older master, order forms that override the master for one purchase, and terms defined by reference to a linked document. Those are the places where commercial terms change, so a repriced term that sits in an amendment gets reported at its old value.
What are page references for in contract extraction?
They tie every field to the page and clause it came from, so a reviewer can check the value without rereading the contract. A useful reference holds the document name, page number, clause number and the quoted text. Page numbers should match the file the reviewer opens, which can differ from the printed page in scanned contracts with cover sheets.
Can the confirmation stage be skipped if the model is good enough?
Not safely. The model is not the limiting factor; it is capable, and capable is not the same as verified. The projects that trusted extraction with no confirmation shipped a wrong date into a live calendar. Make sign off on dates and amounts a required workflow step that users cannot bypass, since the design around the model is what makes the output trustworthy.
How should review of extracted contract data be prioritized?
By confidence and value, so reviewers see the riskiest items first. A workable rule is to check every date and amount in contracts renewing within the next year before anything else, and let clean masters with high confidence scores wait their turn.
How should AI contract extraction accuracy be measured?
On your own messiest contracts, scored field by field and split by document type. Clean masters score well with every tool, so a single figure from a demonstration set tells you little about amendments and order forms, where the errors that cost money appear.
Should unconfirmed fields be indexed for search?
No. Hold them in a review queue until someone signs them off. Once a value appears in search results or a renewal calendar, people treat it as fact, and the confidence score that would have warned them is no longer visible.