1. Understand the Purview Security and Governance System
Map the Purview Landscape and Its Security Boundaries Translate Licensing, Roles, and Governance into an Operating Model
2. Discover, Map, and Curate the Data Estate
Design Data Map Scanning and Metadata Collection Build Unified Catalog, Lineage, Quality, and Data Products
3. Classify and Protect Information
Engineer Sensitive Information Types and Classifiers Design Sensitivity Labels, Publishing, and Auto-Labeling
4. Prevent Unsafe Data Movement
Design DLP Policies from Business Scenarios Extend DLP to Endpoints, Browsers, Teams, and AI
5. Govern the Information Lifecycle
Design Retention Policies and Labels Operate Records, Events, Disposition, and Legal Holds
6. Investigate and Preserve Evidence
Use Purview Audit as Evidence Run eDiscovery Cases, Holds, Searches, and Reviews
7. Manage Human, Communication, and Compliance Risk
Operate Insider Risk and Communication Compliance Responsibly Use Information Barriers and Compliance Manager as Governed Controls
8. Protect Privacy, SharePoint, Microsoft 365, and AI
Secure SharePoint and Microsoft 365 Collaboration Paths Govern Microsoft 365 Copilot and Other Generative AI Protect Privacy and Support Data-Subject Workflows
9. Integrate, Report, and Operate Purview
Integrate Scanners, APIs, Reporting, and Multi-Cloud Sources Run Purview as a Production Security Service Turn DSPM Findings into Data Security Investigations
10. Design and Prove a Complete Purview Program
Build the Purview Target Architecture and Roadmap Capstone: Prove the Security Layer End to End
3. Classify and Protect Information

Engineer Sensitive Information Types and Classifiers

Choose built-in SITs, regex, keyword evidence, EDM, document fingerprinting, and trainable classifiers according to the data and decision.

About this learning content: Courses, lessons, assessments, explanations and illustrations may be created with the help of artificial intelligence. We review and check the material and do our best to avoid incorrect or outdated information, but mistakes, omissions or ambiguous questions may remain. Please verify information before relying on it for professional, security, legal or operational decisions. Read the full notice or report an issue.

In this lesson, you will learn to:

  • Apply a repeatable method for engineer sensitive information types and classifiers in a licensed, governed, and testable Purview environment.

Engineer Sensitive Information Types and Classifiers

This lesson develops a practical understanding of engineer sensitive information types and classifiers and connects design choices to supported capabilities, operational dependencies, user impact, and verifiable evidence.

Detection quality begins with the shape of the evidence

A sensitive information type recognizes evidence in content. Built-in SITs combine primary patterns with supporting evidence, confidence, proximity, and counts. Custom SITs adapt this logic for company formats. Exact Data Match compares detected values with protected fingerprints derived from an authoritative dataset. Document fingerprinting recognizes forms built from a stable template. Trainable classifiers identify meaning across a document rather than a single identifier.

Choose the narrowest evidence that supports the decision. A keyword is useful for context but weak as sole proof. A regex can describe a format but often matches unrelated numbers. EDM is strong for known customer or employee values but requires governed source preparation and refresh. A classifier can identify themes such as source code or business documents, but training examples, licensing, drift, and explainability matter.

The custom SIT false-positive reduction guide and Exact Data Match guide show how to move from broad pattern matching to evidence-backed detection.

Test detections against truth, not convenient samples

Build a labeled evaluation set that includes true examples, close negatives, different formats, languages, old templates, corrupted files, and realistic business context. Keep test material lawful and controlled. Record expected outcome before running the detector so tuning does not redefine success after the fact.

Measure precision and recall in business terms. A detector with high precision may still miss a critical data route. A detector with broad recall may overwhelm users and investigators. Segment results by workload and file type because extraction behavior differs. Review false positives to find missing evidence and false negatives to find unsupported formats or overly strict proximity.

Version the detector and preserve the test set, result, approver, and intended policies. When the source format changes, rerun the evaluation before publishing. For compound keyword requirements, use supporting elements, proximity, and multiple patterns rather than assuming a keyword dictionary provides arbitrary Boolean logic; the Purview keyword AND-logic guide explains the safe patterns.

Resources