datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pii-masking-300k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.Nemotron-PII
Nemotron-PII: Synthesized Data for Privacy-Preserving AI
Dataset Description
Nemotron‑PII is a synthetic, persona‑grounded dataset for training and evaluating detection of Personally Identifiable Information (PII) and Protected Health Information (PHI) in text at production quality. It contains 100,000 English records across 50+ industries with span‑level annotations for 55+ PII/PHI categories, generated with NVIDIA NeMo Data Designer using synthetic personas grounded in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-PII.pii-masking-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Ai4Privacy Community
Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking.
Purpose and Features
Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-200k.pii-masking-openpii-1.5m
OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension)
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Overview
The OpenPII 1.5M dataset extends OpenPII 1M
with a new Asia Pacific corpus, bringing global coverage to 30 languages
across Europe, Americas, and Asia Pacific.
This is the flagship release of the PII-Masking-3M family, the world's
largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1.5m.synthetic_pii_finance_multilingual
Image generated by DALL-E. See prompt for more details
💼 📊 Synthetic Financial Domain Documents with PII Labels
gretelai/synthetic_pii_finance_multilingual is a dataset of full length synthetic financial documents containing Personally Identifiable Information (PII), generated using Gretel Navigator and released under Apache 2.0.
This dataset is designed to assist with the following use cases:
🏷️ Training NER (Named Entity Recognition) models to detect and label PII in… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_pii_finance_multilingual.pii-masking-openpii-1m
OpenPII 1M — Multilingual PII Masking Dataset
Overview
The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types.
Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification pipelines… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1m.gretel-pii-masking-en-v1
Gretel Synthetic Domain-Specific Documents Dataset (English)
This dataset is a synthetically generated collection of documents enriched with Personally Identifiable Information (PII) and Protected Health Information (PHI) entities spanning multiple domains.
Created using Gretel Navigator with mistral-nemo-2407 as the backend model, it is specifically designed for fine-tuning Gliner models.
The dataset contains document passages featuring PII/PHI entities from a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-pii-masking-en-v1.meddies-pii
Meddies PII
Synthetic PII extraction data for multilingual clinical and administrative documents, with language-specific, domain-transfer, translation, and instruction-style views in one Hub repo.
[!IMPORTANT]
This is a synthetic de-identification research artifact for healthcare AI teams. It is not medical advice, not a privacy certification, and not a substitute for task-specific validation on your own data.
If you want to use this dataset in commercial work… See the full description on the dataset page: https://huggingface.co/datasets/Meddies/meddies-pii.pii-masking-400k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-400k.open-pii-masking-500k-ai4privacy
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.pii_benchmark
Russian PII NER Evaluation Dataset
Dataset Description
This dataset is designed for evaluating PII (Personally Identifiable
Information) detection and Named Entity Recognition (NER) systems on
Russian-language text. It targets guardrail and anonymization pipelines that
must reliably find personal data (names, addresses, contacts) and Russian
identity-document numbers (passport, SNILS, INN, OMS, etc.) in text.
The test set combines real, manually annotated examples… See the full description on the dataset page: https://huggingface.co/datasets/redmadrobot-rnd/pii_benchmark.synthetic-pii-function-calling
Dataset Summary
A function calling dataset created by filtering the urchade/synthetic-pii-ner-mistral-v1 dataset.
pile-pii-scrubadub
Dataset Card for pile-pii-scrubadub
Dataset Summary
This dataset contains text from The Pile, annotated based on the personal idenfitiable information (PII) in each sentence.
Each document (row in the dataset) is segmented into sentences, and each sentence is given a score: the percentage of words in it that are classified as PII by Scrubadub.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
This dataset is taken from The… See the full description on the dataset page: https://huggingface.co/datasets/tomekkorbak/pile-pii-scrubadub.pii-masking-65k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
The purpose of the model and dataset is to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs.
The model is a fine-tuned version… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-65k.synthetic_pii_finance_enpii-bench
PII-Bench (ru)
Span-level benchmark for evaluating personal-data (PII) detection in Russian text.
Annotations use explicit character offsets (start, end, type) rather than
IO/BIO/BILOU token tags. This makes the benchmark agnostic to tokenization and
lets you evaluate a complete pipeline — ML model, regular expressions,
post-processing, or a hybrid such as Presidio —
instead of only the model in isolation.
Released alongside GLiNER Guard, a unified safety + PII encoder family:… See the full description on the dataset page: https://huggingface.co/datasets/hivetrace/pii-bench.meddies-pii-mixednym-pii-multilingual-data
nym-pii-multilingual-data
805,000 synthetic, exactly-labeled PII token-classification examples across
~23 languages and 6 scripts, built for training
nym's PII detection models
(e.g. Wismut/nym-pii-multilingual).
Format
JSONL with character-offset spans (offsets index into text as UTF-8 —
compatible with HF fast-tokenizer offset_mapping):
{"text": "Passport Y94316756 issued to Gary Fisher, Ukraine, expires 11/03/2008.",
"entities": [{"start": 9, "end": 18… See the full description on the dataset page: https://huggingface.co/datasets/Wismut/nym-pii-multilingual-data.pii-masking-benchmark
🎭 PIIMB: PII Masking Benchmark
PIIMB measures zero-shot PII masking: a model's ability to mask any PII out-of-the-box, without fine-tuning or label customization.
We designed its evaluation to be character-based and label-agnostic (more details below).
This is an attempt to reflect the most common deployment scenarios where users need broad coverage across document and entity types with no downstream customization (e.g general privacy or data protection).
For specialised use… See the full description on the dataset page: https://huggingface.co/datasets/piimb/pii-masking-benchmark.pii-masking-200k
Purpose and Features
World's largest open source privacy dataset.
The purpose of the dataset is to train models to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs.
The example texts have 54 PII classes (types of sensitive data), targeting 229 discussion subjects / use cases split across business, education, psychology and legal fields, and 5 interactions styles (e.g. casual conversation, formal document, emails… See the full description on the dataset page: https://huggingface.co/datasets/Isotonic/pii-masking-200k.pii-bench
PIIBench
Description
PIIBench is a unified benchmark dataset for PII detection across multiple domains.
Paper
arXiv: http://arxiv.org/abs/2604.15776
Dataset Summary
Total records: 999,940
Entity types: 82
BIO labels: 165 including O
Format: BIO token classification with source text
Structure
Each example contains:
tokens: list of tokens
labels: BIO labels
source: original data source of the sample
text:… See the full description on the dataset page: https://huggingface.co/datasets/Pritesh-2711/pii-bench.pii-bench-zh
PII Bench ZH
Chinese PII (Personally Identifiable Information) detection benchmark dataset. Two subsets covering formal and informal Chinese text, with character-level span annotations.
This is the first open Chinese PII benchmark that covers locale-specific formats (phone, national ID, bank card, license plate, address) with precise offsets.
Disclaimer / 免责声明
This dataset is 100% synthetic and intended solely for research and evaluation purposes. It does not contain any real… See the full description on the dataset page: https://huggingface.co/datasets/wan9yu/pii-bench-zh.ReSID-dataset
Amazon Reviews 2023 (10 Categories, Post-processed)
Paper | Code | Original Source
Overview
This dataset is a curated and post-processed subset of Amazon Reviews 2023. We select 10 product categories and apply a standard preprocessing pipeline widely used in sequential recommendation research. The resulting dataset provides user interaction sequences along with structured item side information.
Categories: 10
Content: user interaction sequences + structured item features… See the full description on the dataset page: https://huggingface.co/datasets/PIIR/ReSID-dataset.kiji-pii-training-data
Kiji PII Detection Training Data
Synthetic multilingual dataset for training PII (Personally Identifiable Information) detection models with token-level entity annotations and coreference resolution.
Dataset Summary
Samples
51,495 (train: 46,345, test: 5,150)
Languages
6 (English, Danish, Dutch, French, Spanish, German)
Countries
20
PII entity types
26
Total entity annotations
397,441 (avg 7.7 per sample)
Coreference clusters
0 (0% of samples)… See the full description on the dataset page: https://huggingface.co/datasets/DataikuNLP/kiji-pii-training-data.nm-kvkk-pii-6K
nm-kvkk-pii-6K
A Turkish dataset for finding personal data in documents — and for working out who each piece of
data belongs to.
5,780 synthetic Turkish documents (contracts, medical records, HR forms, court filings, invoices and
more) annotated with 28,761 personal-data mentions across 118 types, plus 15,535 relations linking
each value to the person it describes. The label set follows the Turkish data-protection law KVKK
(Law No. 6698).
Every document is fabricated. The people… See the full description on the dataset page: https://huggingface.co/datasets/newmindai/nm-kvkk-pii-6K.mapa-eur-lex-pii
Dataset Card for MAPA EUR-LEX PII
A small, multilingual PII-detection benchmark built from the dglover1/mapa-eur-lex test split (itself derived from joelniklaus/mapa). Two English EUR-LEX legal documents were re-annotated by hand with a fine-grained PII scheme, and those labels were then projected onto their professional human translations in 20 other EU languages.
Dataset Details
Dataset Description
The dataset contains 42 documents: two source… See the full description on the dataset page: https://huggingface.co/datasets/piimb/mapa-eur-lex-pii.pii-detection-dataset
Gravitee PII Detection
A harmonized, multi-source corpus for fine-tuning encoder-style PII / NER
models. 25 canonical PII classes, character-level span annotations,
175,881 English examples, 781,052 entity spans.
Published as a single split (train). Hold-out evaluation is expected to be
performed against unrelated external PII corpora rather than against a slice of
this dataset.
Quick start
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/gravitee-io/pii-detection-dataset.synthetic-pii-phi-v2
Synthetic PII/PHI v3
Synthetic clinical/administrative documents in three locales (en-GB, nl-BE,
fr-FR) with span-level PII/PHI annotations, built for training a
multilingual de-identification (token classification) model. Every entity
value — names, IDs, dates, addresses, lab results — is synthetically
generated. No real patient, provider, or organization data appears anywhere
in this dataset; city names and their postcode stems are real (a city name
is not personal data and a… See the full description on the dataset page: https://huggingface.co/datasets/Marawanelbalal/synthetic-pii-phi-v2.maskara-indian-pii-200k
Maskara Indian PII Dataset (Phase 2)
Synthetic Indian PII dataset for training the Maskara NER model. This version has been expanded and diversified for Phase 2 training.
Splits
Split
Rows
Purpose
train
~250,000
Model training
template_disjoint_eval
15,000
Generalization evaluation: entire template families held out
real_world_eval
~2,600
Manually curated real-world Indian text evaluation
Schema
Column
Type
Description… See the full description on the dataset page: https://huggingface.co/datasets/somukandula/maskara-indian-pii-200k.pii-anon
PII-Anon
A CC0 multilingual PII benchmark corpus of 782,677 records carrying 3,107,240 entity annotations across 66 entity types and 60 languages, spanning 7 evaluation dimensions. Each record exposes the five legally-distinct regulatory regime signals (gov-02 / FR-022) as separate reg_* columns — no merged compliance verdict.
Train vs. evaluation substrate
The 159,891 tier3_evaluation records are the EVALUATION substrate of the 782,677-record corpus… See the full description on the dataset page: https://huggingface.co/datasets/subhash-holla/pii-anon.
