You cannot protect data you cannot locate. This guide covers the discovery exercise that establishes where PHI sits across your estate, and the labelling scheme that keeps it protected afterwards. Expect the discovery findings to be uncomfortable and useful in roughly equal measure.
If you have to build it
Sensitive information type selection, auto-labelling policy configuration in simulation mode, label taxonomy design, and the sequencing that avoids breaking clinical workflows on day one.
Why it matters
The problem this solves
Healthcare estates accumulate patient-identifiable data in places nobody planned. Extracts pulled for a project, spreadsheets built for an audit, correspondence attached to a case — each was created for a reason and none was catalogued.
The operational risk is the obvious one. The immediate one, for most providers, is that AI assistance across the estate is now a board-level topic, and it cannot proceed responsibly until you know what is where and who can reach it.
Before you start
Prerequisites
Check these before beginning. Most stalled implementations stall on one of them.
Licensing
Microsoft 365 E5, or E3 with the Compliance or Information Protection add-on. Auto-labelling for SharePoint and OneDrive requires the E5 tier.
Roles
Compliance Administrator or Information Protection Administrator in Purview, plus a privacy officer who can approve the classification scheme.
Governance
An agreed data classification policy. If one does not exist, write it first — a labelling scheme without a policy behind it is arbitrary.
Scope
A decision about which workloads are in scope initially. We recommend SharePoint and OneDrive before Exchange and endpoints.
How it works
The concepts worth understanding first
Configuration is straightforward once these are clear. Skipping them is why most first attempts produce something that works and cannot be maintained.
Sensitive information types are the detection layer
Purview ships with built-in types for health record numbers, insurance identifiers, national identity numbers and similar patterns. These detect content; they do not protect it. Detection accuracy depends on confidence level and instance count, both of which you tune.
Sensitivity labels are the protection layer
A label can apply encryption, restrict access, add visual markings and control whether content can leave the tenant. Crucially, the protection persists with the file, so a labelled document forwarded outside your organization remains protected.
Simulation mode is not optional
Auto-labelling policies run in simulation before enforcement. This tells you what would have been labelled without labelling anything, which is how you discover that a well-intentioned rule would have encrypted every clinical letter in the organization.
Configuration
Step by step
Settings shown are the ones that matter, not every field on the form. Values are starting points to validate against your own environment.
01
Design the label taxonomy first, and keep it small
Four or five labels. More than that and misclassification becomes routine, which is worse than no classification because it produces false confidence.
A workable healthcare scheme is General, Internal, Confidential, and Confidential — PHI, with sublabels only where protection genuinely differs.
Label count
4–5 top level; resist departmental variants
Scope
Files and emails; add Teams and sites only once the file scheme is stable
Encryption
Apply on the PHI label only at first — encryption breaks some downstream tooling
Default label
Set one, so unlabelled content is not the majority state
02
Run discovery before you build any policy
Content Explorer and Activity Explorer show what already matches your sensitive information types across the estate. Do this before creating labels, because the results usually change the taxonomy you were about to build.
Pay particular attention to volume by location. Most PHI concentrates in a small number of sites, and knowing which ones lets you sequence remediation.
Content Explorer
Review by sensitive information type and by location
Confidence level
Start at Medium; High reduces false positives and misses genuine matches
Instance count
Set a minimum — a single matched pattern in a long document is usually noise
03
Create the auto-labelling policy in simulation
Build the rule, scope it to a subset of sites rather than the whole tenant, and run it in simulation for at least a week. Review the matches manually — a sample of a few hundred is enough to see the pattern.
Expect the first simulation to be wrong. That is what it is for.
Policy scope
Start with 2–3 high-volume sites, not the tenant
Mode
Simulation, minimum one week, reviewed before any change
Condition
Sensitive info type plus instance count plus confidence, combined
Exceptions
Exclude template and training libraries explicitly
04
Remediate oversharing before you enforce
Labelling protected content that is already shared with everyone in the organization does not un-share it. Run the oversharing review in parallel — broad-access links, broken inheritance, sites without owners — and fix the highest-risk findings first.
This is the step that most often extends the timeline, and it is the one that actually reduces risk.
05
Enforce, then extend scope gradually
Turn the policy on for the piloted sites. Watch the support queue for a fortnight. Then extend site by site rather than tenant-wide.
Add Exchange and endpoint DLP afterwards, as separate exercises with their own simulation periods.
Verify it worked
Confirm in Content Explorer that labelled volume matches what simulation predicted, within a reasonable margin.
Open a labelled document as a user outside the permitted group and confirm access behaves as designed.
Forward a labelled document to an external address and confirm the protection persists.
Check Activity Explorer for label downgrades — a spike indicates the scheme is fighting a genuine workflow.
Confirm your clinical applications still open labelled documents; encryption occasionally breaks integrations nobody tested.
Best practice
What we do on every engagement of this type
Design the taxonomy before running any policy, then revise it after discovery
Never skip simulation mode, and never shorten it below a full week
Scope initial policies to a few sites rather than the tenant
Run oversharing remediation in parallel with labelling, not after it
Give users a way to report a wrong label, and monitor the reports
Document the classification decisions where your privacy officer will find them
Pitfalls
What catches most first attempts
Every one of these is avoidable, and every one of them is common enough that we check for it by default.
!Too many labels
Every additional label multiplies misclassification. A fifteen-label scheme designed by committee will be applied incorrectly and will produce confident, wrong protection decisions.
!Encrypting before testing integrations
Encryption is the label capability that breaks things. Clinical systems, PDF processors and third-party viewers all handle protected files differently. Test with your actual application set before enforcing.
!Assuming labelling fixes oversharing
It does not. A protected document shared organization-wide is still reachable by everyone in the organization. These are two separate remediation exercises that need to run together.
!Treating discovery as a one-off
Content keeps arriving. Set a recurring review cadence, or the position you established will drift within a year.
Completion checklist
Classification policy agreed and signed off by the privacy officer
Discovery run and reviewed before the taxonomy was finalised
Label taxonomy at five or fewer top-level labels
Auto-labelling policy simulated for a full week and manually sampled
Oversharing remediation completed on the highest-risk sites
Recurring discovery and posture review scheduled
Want a second pair of eyes?
We run scoped discovery against a single workload as a fixed-price engagement, and the findings report is yours whether or not you take the remediation work further with us.