Every eDiscovery vendor now has an AI story. Most are selling the same three things under different names, and almost none of the marketing addresses the question that actually matters to a litigator: if opposing counsel challenges this, can you explain what happened?
This is a practical assessment of where AI genuinely helps in document review, where it does not, and what defensibility requires regardless of the technology.
General information, not legal advice.
First, separate two different technologies
“AI” is doing a lot of work in vendor marketing. Two distinct things are usually bundled together.
Technology-assisted review (TAR) has been in use since roughly 2010. A subject-matter expert codes a training set, a machine-learning classifier extends those judgments across the population, and statistical sampling validates the output. It is well understood, has an established body of case law, and its performance can be measured with recall and precision.
Generative AI — large language models — arrived in review workflows much more recently. Instead of classifying documents against a trained model, it can summarise, answer questions about a corpus, draft privilege log entries, and identify concepts without prior training examples.
These have different track records. TAR has been tested in court for over a decade. Generative AI in review is new enough that the case law is still thin. Treating them as interchangeable, as a lot of marketing does, obscures a real difference in risk.
What AI does well
Prioritisation. Ranking documents by likely responsiveness so reviewers see the important material first. If a matter settles or narrows mid-review, you have already covered the documents that mattered.
Concept clustering. Grouping documents by subject rather than keyword. This surfaces relevant material that keyword searches miss — including documents where people discussed something without using the obvious term. In practice, custodians rarely use the vocabulary that appears in a discovery request.
Email threading and near-duplicate detection. Reviewing a thread once rather than reading the same message quoted eleven times. Unglamorous, and one of the largest genuine cost reductions available.
Summarisation. Condensing long documents so a reviewer can triage faster. Useful for deposition preparation and for getting oriented in an unfamiliar corpus.
First-pass privilege identification. Flagging likely privileged material for attorney review. Note the wording — flagging for review, not deciding.
What AI does not do
It does not make responsiveness calls you can defend without attorney oversight. Responsiveness is a legal judgment about the scope of a specific request in a specific matter. A model can predict what a human coder would likely have decided; it cannot make the judgment itself.
It does not determine privilege. Privilege turns on the relationship between the parties, the purpose of the communication, and whether it has been waived. Models are particularly weak here — an email to in-house counsel discussing a business decision may or may not be privileged, and the distinction rests on facts outside the document.
It does not remove the obligation to validate. Whatever the tool, you need to be able to demonstrate that the process found what it was supposed to find.
It does not make an unclear process defensible. This is the important one.
Defensibility rests on process, not the tool
Courts have not held that TAR is defensible and manual review is not, or that any particular product is approved. What has been examined repeatedly is whether the process was reasonable and whether the party could explain it.
That means four things:
1. A documented protocol. Written before review starts. Who codes the training set, what the coding criteria are, how disagreements are resolved, what the stopping criteria are.
2. Statistical validation. A sample of the documents the system excluded, reviewed by a human, with measured recall. “The system said it was done” is not validation. Numbers you can produce are.
3. Reproducibility. If asked to explain how a specific document ended up excluded, you should be able to answer. This is where some generative AI workflows are genuinely weaker than TAR — a classifier’s decision boundary can be probed and explained; a language model’s reasoning often cannot be reconstructed after the fact.
4. Transparency where appropriate. Many parties disclose their use of TAR and negotiate protocols with opposing counsel. Disclosure is not always required, but a process you would be uncomfortable disclosing is worth reconsidering.
A defensible workflow with a mediocre tool beats an undocumented workflow with the best tool on the market.
The proportionality argument
The strongest practical case for AI-assisted review is not accuracy. It is proportionality.
Federal Rule of Civil Procedure 26(b)(1) limits discovery to what is proportional to the needs of the case. For a matter with a modest amount in controversy, linear review of 300,000 documents is not proportional — the review cost alone can exceed the value of the dispute.
If prioritisation lets a team review 30,000 documents with equivalent recall, that is not a corner cut. It is the mechanism that makes discovery affordable for a case that could not otherwise bear it.
This matters most to smaller firms. Large firms with large matters have always been able to throw reviewers at a problem. Smaller firms have not — and AI-assisted workflows are what put a proportionate process within reach.
Questions to ask any vendor
- Is this TAR, generative AI, or both? If they cannot answer clearly, that is informative.
- How is recall measured, and what will you report to me? You want a number, and a sample you can inspect.
- Can you produce a written protocol before we start?
- If opposing counsel challenges the process, what can you provide? Ask for a specific answer.
- Who reviews the flagged privileged documents? If the answer is “the system handles it,” stop.
- What happens to my data after the matter closes? Hosting costs recur; so does risk.
- Can you explain why a particular document was excluded?
Where this leaves you
AI in document review is real and it works, within limits. Prioritisation, clustering and threading deliver genuine cost reduction. Summarisation saves reviewer time. None of that is hype.
What is hype is the suggestion that any of it removes attorney judgment or the need for a documented, validated process. The technology has changed considerably in three years. What makes a review defensible has not changed at all.
If you are evaluating AI-assisted review for a matter, the useful question is not “how advanced is the model.” It is “what will I be able to show a judge.”
Related
- The Complete eDiscovery Process: A Guide for Law Firms
- eDiscovery & Litigation Support Services
- AI-Driven Data Extraction
Considering AI-assisted review for a matter? Get in touch and we will scope it honestly, including whether it is worth it at your volume.
Pingback: What Does eDiscovery Actually Cost? A Breakdown for Law Firms - EDDM Consulting