Predictive Coding in Microsoft Purview eDiscovery Premium: AI-Powered Document Review

Summarize with:



Written by

— in

ThreatIntelligenceLab.com

Introduction

Predictive coding is the AI feature within Microsoft Purview eDiscovery Premium that transforms how legal teams review documents. Instead of reading 15,000 documents to find the 6,000 that are relevant, you train a model on a small sample, and it prioritises the rest. Reviewers see the most relevant documents first. They stop when relevance drops below a threshold you define. The remaining documents are never reviewed because the model has determined – with measurable confidence – that they are not relevant.

I helped a legal department deploy this for a regulatory investigation last year. They had 15,000 documents collected through a broad keyword search. Manual review would have taken four reviewers three weeks. Predictive coding reduced the review set to 6,000 documents and completed the review in one week. The regulator accepted the methodology because every step was documented and defensible. This guide walks through how to configure predictive coding, train the model, set relevance thresholds, and produce results that stand up to scrutiny.

Predictive coding requires Microsoft 365 E5 or the E5 Compliance add-on with eDiscovery Premium. It is not available in eDiscovery Standard. You also need the eDiscovery Administrator role. If you are new to eDiscovery, start with the eDiscovery Standard guide before moving to Premium and predictive coding.

How Predictive Coding Actually Works – and Why It Is Defensible

Predictive coding is a supervised machine learning model. You show it examples of relevant and non-relevant documents. It learns the patterns that distinguish one from the other. Then it scores every document in the case, ranking them from most to least likely to be relevant. Reviewers start at the top and work down until the relevance rate drops below your threshold.

The defensibility comes from the feedback loop. As reviewers confirm or reject the model’s predictions, those decisions feed back into the model. It continuously improves. By the time reviewers reach the relevance threshold, the model has been trained on hundreds of human-verified examples. The methodology is documented at every step – which documents were used for initial training, how many review rounds were completed, what the relevance scores were at each round, and why the threshold was set where it was.

This is fundamentally different from keyword search. Keyword search returns every document containing your terms, regardless of context. Predictive coding returns documents that are conceptually relevant, even if they use different terminology. A contract dispute might involve documents that never mention “contract” but discuss “agreement,” “terms,” or “scope of work.” Keyword search misses those. Predictive coding finds them because it learns conceptual relevance, not string matching.

Training the Model: The First Review Round Makes or Breaks Accuracy

Predictive coding starts with a training round. You select a sample of documents from the full case set – typically 500 to 1,000 documents – and have experienced reviewers tag each one as relevant or not relevant. This is the seed data. The quality of this initial tagging determines everything that follows.

Use your most experienced reviewers for the training round. Junior reviewers who do not fully understand the case will produce inconsistent tags, and an inconsistently trained model will produce unreliable relevance scores. Each document in the training set should be reviewed by at least two people. Disagreements must be resolved by a third reviewer. This inter-rater reliability is part of what makes the process defensible – you can demonstrate that the model was trained on carefully validated data.

After the training round, the model calculates initial relevance scores for every document in the case. You review the highest-scoring documents and confirm or correct the model’s predictions. Each correction feeds back into the model. After two or three review rounds, the model stabilises – its predictions stop changing significantly because it has seen enough examples. At this point, you set the relevance threshold and begin the final review.

Setting Relevance Thresholds and Measuring Review Efficiency

The relevance threshold is the score below which documents are considered not relevant and are not reviewed. Setting this threshold is a business decision, not a technical one. A lower threshold means more documents reviewed and higher recall – you are less likely to miss something relevant. A higher threshold means fewer documents reviewed and lower cost – but a higher risk of missing relevant material.

For regulatory investigations where missing a relevant document could have legal consequences, I recommend a threshold at the 60th to 70th percentile. This typically captures 90-95% of relevant documents while reducing the review set by 40-50%. For internal investigations where the stakes are lower, a threshold at the 80th percentile can reduce the review set by 60% or more while still capturing the vast majority of relevant material.

Purview tracks two metrics that tell you whether your threshold is working. Recall measures what percentage of all relevant documents the model found. Precision measures what percentage of the documents the model flagged as relevant actually were relevant. Both are visible in the predictive coding dashboard after each review round. If recall drops below 85%, lower your threshold. If precision drops below 70%, your training data may be inconsistent – the model is learning noise rather than signal.

The review efficiency metric that matters most to legal teams is documents reviewed per relevant document found. At the start of review, this number is low – nearly every document is relevant. As reviewers move down the ranked list, this number climbs. When it exceeds your cost-benefit threshold – when reviewers are reading 20 documents to find one relevant one – you stop. This is a cleaner stopping criterion than a fixed relevance score because it directly ties to reviewer effort. Document this metric and the stopping decision for the case record.

Documenting the Process So It Survives a Challenge

The difference between predictive coding that survives a regulatory or legal challenge and predictive coding that does not is documentation. Every decision in the process must be recorded: who selected the training sample, how they selected it, who tagged it, how disagreements were resolved, what the relevance scores were at each round, and why the threshold was set where it was.

Purview automatically captures much of this in the case audit log. The model training history, review round statistics, and relevance score distributions are all stored and exportable. What Purview does not capture is the human reasoning – why reviewers considered a document relevant, why the threshold was chosen, why the review stopped when it did. Create a separate document in the case file that records these decisions alongside the automated metrics.

If the case goes to court or a regulator challenges your process, you need to show that the methodology was reasonable, consistently applied, and documented at every step. Predictive coding has been accepted by courts in multiple jurisdictions – including the UK, US, and EU – precisely because it produces this paper trail. Keyword search often cannot, because there is no record of why specific keywords were chosen and others omitted. Use predictive coding’s transparency as a legal advantage, not just an efficiency tool.

Predictive coding AI document review pipeline showing reduction from 15,000 documents to 6,000 relevant documents with human review feedback loop
Predictive coding uses supervised machine learning to rank documents by relevance. Reviewers see the most relevant documents first and stop when the relevance rate drops below a defensible threshold – typically reducing review volume by 50-60%.

Written by


Comments

Leave a Reply