DLP Test Samples
Download full dataset 2.1 MB
Free · De-identified · CC BY 4.0

Downloadable data samples for testing your DLP

Real-looking sensitive data your DLP should catch, and lookalikes it should ignore. Measure precision and recall across AI agent workflows, GenAI apps, SaaS sharing, email and endpoint exfiltration.

Includes a README and a labels.csv answer key for automatic scoring.

299samples
177positive
122negative
173data types covered
1answer key

Two kinds of samples, two numbers that matter

Most test data only checks whether your DLP fires. Negative samples show whether it fires too often.

Entities and content

Detection has two jobs: find sensitive values inside content, and recognize files that are sensitive as a whole. The dataset tests both.

45document & content samples

Whole files and prompts, judged by what they are

Contracts, payroll, M&A memos and source code; personal bank statements, leases and tax returns; photos of passports and licenses; prompt injections aimed at AI assistants. No single value gives them away.

Browse by data type

Entities, prompts, ID images and whole documents, each with positive and negative samples.

Test every way data leaves

Does your DLP stop sensitive data from leaving, whether a person or an AI agent is moving it?

Score your DLP in three steps

  1. DownloadGet the full dataset. labels.csv lists what should and shouldn't fire for every file.
  2. RunSend each sample through the channel you're testing: a prompt, an agent workflow, an upload, an email, a USB copy.
  3. ScoreRecall = positives caught ÷ all positives. Precision = correct alerts ÷ all alerts. More on scoring →