datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pii-masking-300k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.pii-masking-200k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Ai4Privacy Community
Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking.
Purpose and Features
Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-200k.pii-masking-openpii-1.5m
OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension)
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Overview
The OpenPII 1.5M dataset extends OpenPII 1M
with a new Asia Pacific corpus, bringing global coverage to 30 languages
across Europe, Americas, and Asia Pacific.
This is the flagship release of the PII-Masking-3M family, the world's
largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1.5m.pii-masking-openpii-1m
OpenPII 1M — Multilingual PII Masking Dataset
Overview
The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types.
Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification pipelines… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1m.pii-masking-400k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-400k.open-pii-masking-500k-ai4privacy
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.synthetic-pii-function-calling
Dataset Summary
A function calling dataset created by filtering the urchade/synthetic-pii-ner-mistral-v1 dataset.
pii-masking-65k
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
Purpose and Features
The purpose of the model and dataset is to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs.
The model is a fine-tuned version… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-65k.nym-pii-multilingual-data
nym-pii-multilingual-data
805,000 synthetic, exactly-labeled PII token-classification examples across
~23 languages and 6 scripts, built for training
nym's PII detection models
(e.g. Wismut/nym-pii-multilingual).
Format
JSONL with character-offset spans (offsets index into text as UTF-8 —
compatible with HF fast-tokenizer offset_mapping):
{"text": "Passport Y94316756 issued to Gary Fisher, Ukraine, expires 11/03/2008.",
"entities": [{"start": 9, "end": 18… See the full description on the dataset page: https://huggingface.co/datasets/Wismut/nym-pii-multilingual-data.pii-masking-benchmark
🎭 PIIMB: PII Masking Benchmark
PIIMB measures zero-shot PII masking: a model's ability to mask any PII out-of-the-box, without fine-tuning or label customization.
We designed its evaluation to be character-based and label-agnostic (more details below).
This is an attempt to reflect the most common deployment scenarios where users need broad coverage across document and entity types with no downstream customization (e.g general privacy or data protection).
For specialised use… See the full description on the dataset page: https://huggingface.co/datasets/piimb/pii-masking-benchmark.pii-bench
PIIBench
Description
PIIBench is a unified benchmark dataset for PII detection across multiple domains.
Paper
arXiv: http://arxiv.org/abs/2604.15776
Dataset Summary
Total records: 999,940
Entity types: 82
BIO labels: 165 including O
Format: BIO token classification with source text
Structure
Each example contains:
tokens: list of tokens
labels: BIO labels
source: original data source of the sample
text:… See the full description on the dataset page: https://huggingface.co/datasets/Pritesh-2711/pii-bench.pii-bench-zh
PII Bench ZH
Chinese PII (Personally Identifiable Information) detection benchmark dataset. Two subsets covering formal and informal Chinese text, with character-level span annotations.
This is the first open Chinese PII benchmark that covers locale-specific formats (phone, national ID, bank card, license plate, address) with precise offsets.
Disclaimer / 免责声明
This dataset is 100% synthetic and intended solely for research and evaluation purposes. It does not contain any real… See the full description on the dataset page: https://huggingface.co/datasets/wan9yu/pii-bench-zh.nm-kvkk-pii-6K
nm-kvkk-pii-6K
A Turkish dataset for finding personal data in documents — and for working out who each piece of
data belongs to.
5,780 synthetic Turkish documents (contracts, medical records, HR forms, court filings, invoices and
more) annotated with 28,761 personal-data mentions across 118 types, plus 15,535 relations linking
each value to the person it describes. The label set follows the Turkish data-protection law KVKK
(Law No. 6698).
Every document is fabricated. The people… See the full description on the dataset page: https://huggingface.co/datasets/newmindai/nm-kvkk-pii-6K.mapa-eur-lex-pii
Dataset Card for MAPA EUR-LEX PII
A small, multilingual PII-detection benchmark built from the dglover1/mapa-eur-lex test split (itself derived from joelniklaus/mapa). Two English EUR-LEX legal documents were re-annotated by hand with a fine-grained PII scheme, and those labels were then projected onto their professional human translations in 20 other EU languages.
Dataset Details
Dataset Description
The dataset contains 42 documents: two source… See the full description on the dataset page: https://huggingface.co/datasets/piimb/mapa-eur-lex-pii.korean-pii-dataset
Korean Synthetic PII Dataset
한국어 문장 안의 개인정보(PII)를 탐색·분석하거나 토큰 분류 모델을 학습할 수 있도록 제작한 합성 데이터셋입니다. 모든 이름·번호·주소와 문장은 합성 값이며 실제 개인의 개인정보를 의도적으로 포함하지 않았습니다.
라이선스
이 데이터셋은 Creative Commons Attribution 4.0 International (CC BY 4.0)으로 제공합니다. 재배포·수정·상업적 이용이 가능하지만, townboy/korean-pii-dataset과 원 저작자를 표시해야 합니다. 라이선스 전문은 저장소의 LICENSE 파일을 확인하세요.
구성
항목
수량
전체 문서
11,732
Train
9,227
Validation
1,510
Test
995
PII span
55,627
PII 유형
33
BIO 라벨
67 (O… See the full description on the dataset page: https://huggingface.co/datasets/townboy/korean-pii-dataset.pii-masking-health-phi-preview
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
PII Masking Personal Health & Medical Information (PHI) — Preview
50 sample entries from the PII-Masking-2M European release by AI4Privacy.
Source text and PII values are redacted in this preview. Contact us… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-health-phi-preview.k-pii-bench
K-PII-Bench
A Benchmark for Korean Personal Information Detection.
12 domains · 18 PII types (37 BIO labels) · ~300,000 synthetic Korean documents · skeleton-disjoint splits.
Code & full docs: https://github.com/woohyun212/k-pii-bench
License: data CC BY 4.0 · code Apache-2.0
Paper: K-PII-Bench: A Benchmark for Korean Personal Information Detection (Language Resources and Evaluation, Springer Nature)
Load
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/woohyun212/k-pii-bench.rizzo-pii-it-dataset
rizzo-pii · Italian PII dataset 🦔🛡️
The training & validation data behind rizzoaiacademy/rizzo-pii-0.3B
— a PII token-classification model for Italian legal text, covering 22 categories of
personal data including the Italian legal identifiers (codice fiscale, partita IVA,
dati catastali) that no other open PII model handles.
Everything here serves one goal: anonymize legal documents locally before sending them to a
closed LLM (anonymize → reversible local dictionary → API →… See the full description on the dataset page: https://huggingface.co/datasets/rizzoaiacademy/rizzo-pii-it-dataset.bc-pii-ko
bc-pii-ko — 한국어 개인정보(PII) 탐지 학습 데이터셋
ai4privacy/pii-masking-openpii-1.5m
의 한국어 서브셋을 BC 개인정보 정책 유형 체계로 가공하고 도메인 증강을 더한
토큰 분류(NER) 학습용 데이터셋. 전량 합성/공개 데이터 — 실존 개인정보 미포함
(주민번호 등은 형식 규칙만 준수한 난수, 카드번호는 테스트 IIN + Luhn).
Splits
split
행 수
성격
train
62,194
원본 가공 19,158 + 값치환 사본 38,316 + 절조합 주입 4,000 + 네거티브 720
validation
5,273
증강 미포함 원본 분포 (튜닝·에폭 선택용)
test
2,067
증강 미포함 원본 분포 — 대표 지표는 여기서
test_gap
640
학습에 없는 절로만 조합한 갭 유형(계좌·CI·IP 등) 일반화 진단 전용
규약: train 만… See the full description on the dataset page: https://huggingface.co/datasets/sehe1121/bc-pii-ko.pii-masking-mini-10k
PII Masking Mini: Multilingual Sample
A mini-sized stratified sample of pii-masking-openpii-1.5m,
the flagship release of the PII-Masking-3M family. Sampled proportionally by
(source_dataset, language) so every locale and label gets representation.
Asia Pacific rows appear first.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Dataset Details
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-mini-10k.PI-Indic-Align
PI-Indic-Align: Persona-Instruction Alignment for Indian Languages
🦚 PI-Indic-Align: Persona-Instruction Alignment for Indian Languages
Teaching AI to speak the languages of India, one persona at a time!
Dataset Description
PI-Indic-Align is a large-scale benchmark dataset for evaluating persona-instruction alignment across 12 major Indian languages. The dataset contains 600,000 culturally grounded persona-instruction pairs (50,000 per language) designed… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/PI-Indic-Align.pii-masking-micro-100k
PII Masking Micro: Multilingual Sample
A micro-sized stratified sample of pii-masking-openpii-1.5m,
the flagship release of the PII-Masking-3M family. Sampled proportionally by
(source_dataset, language) so every locale and label gets representation.
Asia Pacific rows appear first.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Dataset Details
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-micro-100k.french-pii-eval
French PII Evaluation Dataset
A curated French PII detection evaluation and training dataset, built for benchmarking TheoDB/privacy-filter-fr.
Dataset Structure
Split
Examples
Purpose
test.jsonl
2,500
Held-out evaluation — never used in training
test_english.jsonl
426
English regression check
train.jsonl
57,248
Training data
val.jsonl
500
Validation data
Label Taxonomy
8 PII classes (same as openai/privacy-filter):
Class… See the full description on the dataset page: https://huggingface.co/datasets/TheoDB/french-pii-eval.pii-masking-financial-pfi-preview
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
PII Masking Personal Financial Information (PFI) — Preview
50 sample entries from the PII-Masking-2M European release by AI4Privacy.
Source text and PII values are redacted in this preview. Contact us for full… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-financial-pfi-preview.NUS-LT-PII-corpus
NUS Lithuanian PII Corpus
Description
Lithuanian text annotated for personal information (PII), spanning three subject domains — administrative, scientific, and media — plus a stratified validation set. The corpus covers 24 entity types: 16 general categories (PER, LOC, ORG, …) and 8 GDPR special-category "sensitive" entities (REL, POL, SEX, GENDER, MAR, FAM, ETH, HEALTH).
Dataset Summary
Subsets: 4 (3 training categories + 1 validation set)
Total records: 41… See the full description on the dataset page: https://huggingface.co/datasets/VytautoDidziojoUniversitetas/NUS-LT-PII-corpus.pii-synthetic-ru
PII Synthetic Dataset (Russian) — pii-synthetic-ru
Синтетический датасет для обучения NER-детектора персональных данных на русском языке.
Содержит 4 500 примеров с аннотациями сущностей NAME (ФИО) и ADDRESS (адрес),
а также негативные примеры без ПД.
Создан в рамках проекта PIIDetector — гибридного
детектора ПД для русского языка на базе Microsoft Presidio.
Статистика
Метрика
Значение
Всего примеров
4 500
Только NAME
1 113
Только ADDRESS
952
NAME… See the full description on the dataset page: https://huggingface.co/datasets/alrosait/pii-synthetic-ru.pii-masking-nano-1k
PII Masking Nano: Multilingual Sample
A nano-sized stratified sample of pii-masking-openpii-1.5m,
the flagship release of the PII-Masking-3M family. Sampled proportionally by
(source_dataset, language) so every locale and label gets representation.
Asia Pacific rows appear first.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Dataset Details
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-nano-1k.ko-pii-ner-100k
한국 PII 특화 학습용 데이터셋 (ko_pii_v1)
한국 고유 식별자와 조사 결합 경계를 정면으로 다루는 한국어 PII NER 학습 데이터셋.
1. 개요
학습용 98,845건 + 외부 평가용 홀드아웃 2,006건
라벨 20종 3티어 / BIO 41 클래스
시드 42, 검증자릿수 정책 invalid, 사용 티어 [1, 2, 3]
포맷: JSONL. {id, text, spans:[{start,end,label,value}], meta}
이 레포의 NERPreprocessor span 포맷과 동일해 학습 경로 수정 없이 사용 가능
1-1. 이 데이터셋으로 학습한 모델
seongyeon1/ko-pii-ner-roberta-base
(klue/roberta-base 파인튜닝, CC-BY-SA-4.0)
학습에 한 번도 쓰이지 않은 KDPII 공식 test split 기준 F5 0.9344.
내부… See the full description on the dataset page: https://huggingface.co/datasets/seongyeon1/ko-pii-ner-100k.pii-intent-detection-multilingual
PII Intent Detection - Multilingual Dataset (TR/AR/EN)
A multilingual dataset for training PII (Personally Identifiable Information) sharing intent classifiers. Covers Turkish, Arabic, and English with 41,427 labeled samples across 9 entity types.
Dataset Description
This dataset was created for content moderation on creator-brand collaboration platforms. The goal is to detect whether a user intends to share personal contact information to move communication off-platform… See the full description on the dataset page: https://huggingface.co/datasets/gorkem371/pii-intent-detection-multilingual.privy
Dataset Card for "privy-english"
Dataset Summary
A synthetic PII dataset generated using Privy, a tool which parses OpenAPI specifications and generates synthetic request payloads, searching for keywords in API schema definitions to select appropriate data providers. Generated API payloads are converted to various protocol trace formats like JSON and SQL to approximate the data developers might encounter while debugging applications.
This labelled PII dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/piimb/privy.
