Electronic discovery is where a large share of modern litigation budget is won or lost. The rules are not complicated, but the decisions made in the first two weeks of a matter determine what the next eighteen months cost.
This guide walks through the eDiscovery process stage by stage: what happens, what it costs, where matters go wrong, and what to decide early. It is written for attorneys and legal operations staff who need a working understanding of the process — not a technical manual.
General information, not legal advice. Requirements vary by jurisdiction and matter.
What Is eDiscovery?
eDiscovery is the process of identifying, preserving, collecting, processing, reviewing and producing electronically stored information (ESI) in response to litigation, a regulatory request or an internal investigation.
The scope is wider than most people expect. ESI includes email, documents on file shares and personal drives, messages in Slack or Teams, text messages and mobile data, database records, cloud applications, voicemail, calendar entries and metadata attached to all of it.
It is the metadata that often matters most. When a document was created, who modified it and when, who received it — that context frequently carries more evidentiary weight than the document text.
What Is the EDRM Model?
The Electronic Discovery Reference Model (EDRM) is the framework the industry uses to describe the eDiscovery lifecycle. It defines nine stages:
- Information Governance — how an organisation manages records before litigation arises
- Identification — determining what data exists and where
- Preservation — ensuring relevant data is not altered or destroyed
- Collection — gathering data defensibly
- Processing — reducing and normalising data for review
- Review — assessing documents for responsiveness and privilege
- Analysis — evaluating content for patterns, key facts and case strategy
- Production — delivering documents to the requesting party
- Presentation — displaying material at deposition, hearing or trial
Two things are commonly misunderstood.
It is a reference model, not a sequence. Matters move backwards and forwards between stages constantly. New custodians surface during review. Production format disputes send you back to processing.
Volume drops sharply as you move right. A collection of one terabyte might reduce to 200 gigabytes after processing and 20 gigabytes after culling. The cost of each stage rises as volume falls — which is exactly why work done early pays for itself.
Stage 1: Identification
Identification means answering two questions: who are the custodians, and where does their data live?
Custodians are the people whose records are likely relevant. Data sources include email servers, file shares, laptops, mobile devices, collaboration platforms, cloud storage, databases and increasingly SaaS applications that IT may not fully inventory.
The common failure here is scoping by job title rather than by involvement. The assistant who scheduled the meetings often holds more relevant material than the executive who attended them.
Practical step: interview custodians directly. Ask what systems they actually use, not what they are supposed to use. Shadow IT — personal drives, unofficial group chats, personal email for work — is where preservation failures originate.
Stage 2: Preservation
Once litigation is reasonably anticipated, the duty to preserve attaches. That is earlier than most people assume — it is not when the complaint is filed.
Preservation means suspending routine deletion and ensuring relevant ESI is not modified. In practice:
- Issue a written litigation hold to identified custodians
- Suspend auto-delete policies on email and collaboration tools
- Notify IT to halt recycling of backup media
- Document everything — who was notified, when, what they were told, and their acknowledgement
- Reissue reminders periodically; a hold issued once and never mentioned again is weak
Under Federal Rule of Civil Procedure 37(e), a party that fails to take reasonable steps to preserve ESI can face curative measures, adverse inference instructions, or in cases of intent to deprive, terminating sanctions.
The documentation matters as much as the act. If you cannot show what you did, you may as well not have done it.
Stage 3: Collection
Collection is the defensible capture of data from identified sources.
“Defensible” means three things: metadata is preserved, chain of custody is documented, and the process can be explained and repeated by a competent third party.
This is why self-collection is risky. When a custodian drags files into a folder and emails them over, creation and modification dates change, folder structure is lost, and there is no record of what was not collected. Courts have taken a dim view of unsupervised custodian self-collection.
Targeted collection versus full forensic imaging. A full forensic image captures everything including deleted-file remnants — appropriate where spoliation is alleged or the device itself is evidence. Targeted collection captures defined sources and date ranges, and is proportionate for the majority of civil matters. Over-collecting is the most common and most expensive early mistake, because every gigabyte collected is processed, hosted and potentially reviewed.
Stage 4: Assessment and Early Case Evaluation
Before committing to review, find out what you actually have.
Assessment examines the collected set: total volume, document counts, date distribution, custodian distribution, file types, duplication rate and how much is non-substantive system material.
This is where scope negotiations should be grounded. Going into a Rule 26(f) conference knowing your collection is 400 gigabytes across nine custodians, 60% duplicative, with a date range extending two years past the relevant period, gives you concrete arguments for narrowing. Going in without that leaves you agreeing to terms you cannot cost.
Stage 5: Processing
Processing converts collected data into a reviewable form and — critically — reduces its volume.
Standard steps:
- De-duplication — identical documents appear many times across custodians; deduplicating removes redundant review
- De-NISTing — removing known system and application files using the NIST reference library
- Date filtering — excluding material outside the relevant period
- Search term application — applying agreed terms to isolate likely-responsive material
- Email threading — grouping conversations so a reviewer reads the thread once rather than every message
- Text extraction and OCR — making scanned material searchable
- Exception handling — flagging corrupt, encrypted or password-protected files rather than silently dropping them
That last point deserves attention. Exception reports are routinely ignored, and an encrypted archive that never got processed is exactly what surfaces at the worst moment.
This stage produces the largest cost reduction in the entire process. Culling before review is far cheaper than reviewing and then discarding.
Stage 6: Review
Review is where documents are assessed for responsiveness, privilege and issue coding.
It is almost always the single largest cost in a matter, because it scales with document count and requires human judgment.
Linear review means reading everything. Predictable, thorough, expensive, and impractical above modest volumes.
Technology-assisted review (TAR) uses machine learning to prioritise or classify. A subject-matter expert codes a training set, the system extends those judgments across the population, and sampling validates the result. Courts have accepted TAR workflows since 2012.
Privilege review runs alongside. A Rule 502(d) order is worth securing early — it protects against subject-matter waiver if privileged material is inadvertently produced, and costs nothing but the asking.
AI and Document Review: What Attorneys Should Know
Generative AI is now being applied to review, and the marketing has run ahead of the practice. A realistic assessment:
What it does well. Prioritising likely-responsive documents, clustering by concept, summarising long documents, surfacing themes across a corpus, and drafting first-pass privilege logs.
What it does not do. Replace attorney judgment on responsiveness calls, make privilege determinations you can defend without review, or remove the need for a documented, validated process.
Defensibility rests on process, not the tool. Regardless of what technology you use, you need a written protocol, a defined training and validation approach, sampling with measurable results, and the ability to explain and reproduce what happened. A workflow you cannot explain to a judge is not defensible, however sophisticated.
The practical benefit is proportionality. If AI-assisted prioritisation lets a team review 30,000 documents instead of 300,000 with equivalent recall, that is a real reduction in cost and calendar time — which matters most to firms without unlimited review budgets.
Stage 7: Production
Production is delivery to the requesting party.
Format should be agreed early, ideally at the Rule 26(f) conference. Under Rule 34, if the request does not specify a form, production must be in the form in which the material is ordinarily maintained or a reasonably usable form.
Common formats:
- Native — original files with metadata intact; standard for spreadsheets, where converting destroys formulas
- TIFF or PDF with load file — images plus a separate file carrying metadata and text
- Hybrid — natives for spreadsheets and media, images for everything else
Production also involves Bates numbering, applying confidentiality designations under the protective order, redaction, and a privilege log for withheld material.
Format disputes discovered after production are expensive. Re-producing a set in a different format means reprocessing, re-numbering and re-checking. Settle it in writing before the first production goes out.
What Drives eDiscovery Cost
Pricing is usually quoted per gigabyte, but the per-GB rate is a poor predictor of total spend. Costs distribute across four areas:
- Collection — typically per custodian or device; one-time
- Processing — per gigabyte ingested; one-time
- Hosting — per gigabyte per month, recurring for the life of the matter
- Review — per document or per hour, usually the largest line item
Two implications follow.
Hosting recurs. A matter that stays live for two years pays hosting twenty-four times. Releasing or archiving data when it is no longer needed is a real saving that is routinely forgotten.
Review scales with document count, not gigabytes. Anything that reduces the document population — deduplication, threading, targeted collection, aggressive culling — reduces the largest cost.
A lower per-GB rate applied to an over-collected dataset costs more than a higher rate applied to a properly scoped one.
An eDiscovery Checklist for Litigation
At the outset of a matter:
- Determine when the duty to preserve attached
- Issue a written litigation hold; document distribution and acknowledgements
- Suspend auto-deletion and backup recycling
- Interview custodians about the systems they actually use
- Inventory data sources, including cloud and mobile
- Assess volume before agreeing scope
- Agree search terms, date ranges and custodians with opposing counsel where possible
- Agree production format in writing
- Secure a Rule 502(d) order
- Confirm who bears which costs
- Keep a written record of every scoping decision and why it was made
Most discovery disputes trace back to gaps in preservation and documentation — not to review errors.
eDiscovery for Small and Mid-Size Firms
Most of this industry is built for AmLaw 100 firms and enterprise matters. Smaller firms are often quoted enterprise pricing for a matter that does not need enterprise infrastructure, and routed through an account manager rather than the people handling the data.
A smaller matter does not need a smaller version of an enterprise process. It needs a proportionate one: targeted collection instead of blanket imaging, aggressive early culling, review technology used where volume justifies it, and honest scoping before anyone commits.
If you are handling your first eDiscovery matter, the most valuable thing you can do is have the scoping conversation before the collection, not after.
How EDDM Consulting Can Help
We support the full lifecycle — collection, assessment, processing, review and production — alongside document management, court reporting and certified legal translation.
We work with firms of every size, and you deal directly with the people handling your data.
Get in touch to scope a matter.