Real-looking sensitive data your DLP should catch, and lookalikes it should ignore. Measure precision and recall across AI agent workflows, GenAI apps, SaaS sharing, email and endpoint exfiltration.
Includes a README and a labels.csv answer key for automatic scoring.
Positive samplePrompt to an AI assistant
Summarize this visit note for the referral letter: Patient Maria Chen (DOB 03/14/1968) was seen for type 2 diabetes. A1C is 8.2%; metformin increased to 1000 mg.
BlockedProtected health information: patient + diagnosis
Negative samplePrompt to an AI assistant
Summarize these meeting notes: Bob Jones met with Summit Paving on 03/15/2026 about resurfacing the parking lot at Denver Cancer Center.
AllowedName, date and hospital, but no patient care
299samples
177positive
122negative
173data types covered
1answer key
Two kinds of samples, two numbers that matter
Most test data only checks whether your DLP fires. Negative samples show whether it fires too often.
Positive samples
Measure recall
Sensitive data in the context where it really appears: support tickets, source code, spreadsheets, screenshots, ID photos and business documents. Your DLP should flag every one.
Lookalikes that aren't sensitive: card-shaped transaction IDs, key-shaped config values, medical words with no patient data. Use these to judge the noise in your DLP.
Detection has two jobs: find sensitive values inside content, and recognize files that are sensitive as a whole. The dataset tests both.
254entity samples
Sensitive values in context
SSNs, card numbers, API keys, passwords and patient data, embedded where they really appear: support tickets, source code, CSV exports, screenshots and chat messages.
Contracts, payroll, M&A memos and source code; personal bank statements, leases and tax returns; photos of passports and licenses; prompt injections aimed at AI assistants. No single value gives them away.