Engineer Sensitive Information Types and Classifiers
Choose built-in SITs, regex, keyword evidence, EDM, document fingerprinting, and trainable classifiers according to the data and decision.
In this lesson, you will learn to:
- Apply a repeatable method for engineer sensitive information types and classifiers in a licensed, governed, and testable Purview environment.
Engineer Sensitive Information Types and Classifiers
This lesson develops a practical understanding of engineer sensitive information types and classifiers and connects design choices to supported capabilities, operational dependencies, user impact, and verifiable evidence.
Detection quality begins with the shape of the evidence
A sensitive information type recognizes evidence in content. Built-in SITs combine primary patterns with supporting evidence, confidence, proximity, and counts. Custom SITs adapt this logic for company formats. Exact Data Match compares detected values with protected fingerprints derived from an authoritative dataset. Document fingerprinting recognizes forms built from a stable template. Trainable classifiers identify meaning across a document rather than a single identifier.
Choose the narrowest evidence that supports the decision. A keyword is useful for context but weak as sole proof. A regex can describe a format but often matches unrelated numbers. EDM is strong for known customer or employee values but requires governed source preparation and refresh. A classifier can identify themes such as source code or business documents, but training examples, licensing, drift, and explainability matter.
The custom SIT false-positive reduction guide and Exact Data Match guide show how to move from broad pattern matching to evidence-backed detection.
Test detections against truth, not convenient samples
Build a labeled evaluation set that includes true examples, close negatives, different formats, languages, old templates, corrupted files, and realistic business context. Keep test material lawful and controlled. Record expected outcome before running the detector so tuning does not redefine success after the fact.
Measure precision and recall in business terms. A detector with high precision may still miss a critical data route. A detector with broad recall may overwhelm users and investigators. Segment results by workload and file type because extraction behavior differs. Review false positives to find missing evidence and false negatives to find unsupported formats or overly strict proximity.
Version the detector and preserve the test set, result, approver, and intended policies. When the source format changes, rerun the evaluation before publishing. For compound keyword requirements, use supporting elements, proximity, and multiple patterns rather than assuming a keyword dictionary provides arbitrary Boolean logic; the Purview keyword AND-logic guide explains the safe patterns.
Resources
- Microsoft Learn reference for Engineer Sensitive Information Types and Classifiers — Official Microsoft documentation supporting the capability, prerequisites, and current product behavior taught in this lesson.