DocPII contains 1101 high-quality document samples enriched with embedded personally identifiable information (PII). Designed to evaluate context-aware redaction systems, it provides realistic, full-document contexts—a notable advancement over sentence-level datasets.
All documents have been manually reviewed for accuracy, coherence, and redaction alignment, ensuring data quality for benchmarking and development.
DocPII extends the Gretel PII Masking dataset by embedding its PII entities into longer, domain-specific documents. These were generated using GPT-4.1-nano with prompt engineering to simulate authentic formats across healthcare, finance, and other sectors.
This extension effort enables more rigorous evaluation of document-level redaction, information extraction, and privacy protection systems.
uid (string): Unique identifiertext (string): Full synthetic document with embedded PIIentities (list):
entity (string): Sensitive entity valuetypes (array of strings): PII categories (e.g., NAME, PHONE_NUMBER)redaction_query (string): Natural language instructiondomain (string): General domain (e.g., healthcare)document_type (string): Specific document type (e.g., tax form)document_description (string): Summary of document's function and formatentity_count (integer): Total number of embedded PII entitiesMost public redaction datasets lack full-document context. DocPII addresses this by offering realistic workflows and instructions embedded in industry-style documents, better reflecting actual use cases in regulated sectors.
All entries were manually reviewed to validate coherence, instruction alignment, and redaction relevance.
DocPII includes a wide variety of synthetic PII types, such as:
No real-world or user data is included.
DocPII supports the development of safer, privacy-respecting AI by enabling rigorous evaluation of redaction and PII detection systems used in sensitive domains.
Potential sources of bias include:
If you use this dataset in academic or commercial work, please cite it as:
@dataset{nutrientio_2024_docpii,
title = {DocPII: Contextual Redaction Benchmark Dataset},
author = {Nutrient.io},
year = {2025},
howpublished = {\url{https://huggingface.co/datasets/nutrientdocs/synthetic_labeled_redaction_instruction_en_v1}},
note = {Synthetic dataset for document-contextual PII redaction evaluation, based on gretelai/gretel-pii-masking-en-v1}
}
3 commits
DocPII contains 1101 high-quality document samples enriched with embedded personally identifiable information (PII). Designed to evaluate context-aware redaction systems, it provides realistic, full-document contexts—a notable advancement over sentence-level datasets.
All documents have been manually reviewed for accuracy, coherence, and redaction alignment, ensuring data quality for benchmarking and development.
DocPII extends the Gretel PII Masking dataset by embedding its PII entities into longer, domain-specific documents. These were generated using GPT-4.1-nano with prompt engineering to simulate authentic formats across healthcare, finance, and other sectors.
This extension effort enables more rigorous evaluation of document-level redaction, information extraction, and privacy protection systems.
uid (string): Unique identifiertext (string): Full synthetic document with embedded PIIentities (list):
entity (string): Sensitive entity valuetypes (array of strings): PII categories (e.g., NAME, PHONE_NUMBER)redaction_query (string): Natural language instructiondomain (string): General domain (e.g., healthcare)document_type (string): Specific document type (e.g., tax form)document_description (string): Summary of document's function and formatentity_count (integer): Total number of embedded PII entitiesMost public redaction datasets lack full-document context. DocPII addresses this by offering realistic workflows and instructions embedded in industry-style documents, better reflecting actual use cases in regulated sectors.
All entries were manually reviewed to validate coherence, instruction alignment, and redaction relevance.
DocPII includes a wide variety of synthetic PII types, such as:
No real-world or user data is included.
DocPII supports the development of safer, privacy-respecting AI by enabling rigorous evaluation of redaction and PII detection systems used in sensitive domains.
Potential sources of bias include:
If you use this dataset in academic or commercial work, please cite it as:
@dataset{nutrientio_2024_docpii,
title = {DocPII: Contextual Redaction Benchmark Dataset},
author = {Nutrient.io},
year = {2025},
howpublished = {\url{https://huggingface.co/datasets/nutrientdocs/synthetic_labeled_redaction_instruction_en_v1}},
note = {Synthetic dataset for document-contextual PII redaction evaluation, based on gretelai/gretel-pii-masking-en-v1}
}
3 commits