You cannot protect an unknown asset
Data protection begins with visibility. If an organization cannot name its sensitive data, its approved locations, or its accountable owners, a DLP policy has to guess. Guessing usually creates one of two outcomes: weak rules that miss real leakage, or broad rules that interrupt ordinary work.
The first objective is not a perfect inventory of every file. It is a reliable view of the data that could create meaningful harm if exposed.
Build a minimum viable data inventory
Start with high-value business processes and record:
- Data set: What information is involved?
- Business owner: Who decides how it may be used and shared?
- Systems and locations: Where is the authoritative copy, and where do working copies appear?
- People and roles: Who requires access?
- External recipients: Which processors, partners, or customers legitimately receive it?
- Retention need: How long is it required?
- Impact: What happens if it is disclosed, altered, or unavailable?
- Required controls: Which sharing, encryption, access, or monitoring rules apply?
Interview people who perform the work. Technical administrators can identify repositories, but process owners know why exports are created, which exceptions are real, and when a third party must receive the information.
Use a small classification model
A classification scheme should help a person make a decision in seconds. Four levels are often enough:
- Public: Approved for anyone to see.
- Internal: Routine business information intended for the workforce and approved collaborators.
- Confidential: Sensitive business or personal data that needs controlled sharing.
- Highly confidential: Information where exposure could cause severe harm and access should be tightly limited.
The names matter less than the handling rules behind them. For each level, define whether external sharing is allowed, whether encryption is required, which storage locations are approved, and whether users can override a restriction.
Avoid labels that depend on legal expertise. A user should not need to interpret several regulations before deciding how to handle a document.
Classification needs more than labels
Classification can come from several signals:
- Manual judgment: A person applies a label because they understand the document's context.
- Pattern matching: A tool recognizes structured values such as account numbers or identifiers.
- Exact matching: Known records are detected with higher precision against an approved reference set.
- Document structure: A standard form or template is recognized.
- Content-based classification: A classifier identifies meaning that cannot be expressed as a simple pattern.
- Context: The repository, owner, recipient, device, or activity changes the risk.
No single method is sufficient. Pattern detection may find an identifier but cannot always tell whether the document is a test file, a public form, or a high-risk customer export. Context and volume can change the decision.
Most platforms combine these methods differently. They may use different names, detectors, confidence models, and enforcement actions. Learn the classification method first, then verify how the platform you operate implements it, where it is supported, and how its results can feed policy, sharing, retention, or DLP controls.
Apply this model
Once the data and classification approach are clear, continue with the guidance for your environment:
- Microsoft 365 guidance explains Purview classification, labels, detectors, and DLP controls.
- Google Workspace guidance explains Google-specific locations, rules, sharing controls, and evidence.
- Sekurzen Guard guidance explains Guard detectors, workload actions, operational visibility, and coexistence questions.
Define ownership before automation
Security teams can operate the platform, but they should not decide every acceptable use of business data. Use a simple responsibility model:
- the business owner defines sensitivity and legitimate use
- the data or privacy function translates regulatory and contractual needs
- the security team designs and monitors controls
- the platform team implements and maintains technical configuration
- the service desk and incident team handle user questions and events
Each high-impact policy needs a named decision maker for exceptions. Without one, administrators either block legitimate work or quietly weaken the rule.
Validate with representative content
Before automation, assemble a controlled test set containing:
- true examples that the rule should detect
- similar but non-sensitive examples that should not match
- edge cases with different languages, formats, or volumes
- encrypted, scanned, compressed, and image-based files where relevant
Measure false positives and false negatives separately. A rule that detects 100 percent of test items but interrupts thousands of harmless messages is not ready for enforcement.
The next part maps the channels through which known sensitive data can leak.

