datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pii-masking-openpii-1.5m
OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension)
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Overview
The OpenPII 1.5M dataset extends OpenPII 1M
with a new Asia Pacific corpus, bringing global coverage to 30 languages
across Europe, Americas, and Asia Pacific.
This is the flagship release of the PII-Masking-3M family, the world's
largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1.5m.pii-masking-openpii-1m
OpenPII 1M — Multilingual PII Masking Dataset
Overview
The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types.
Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification pipelines… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1m.open-pii-masking-500k-ai4privacy
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.openpii-masking-nano-1k
OpenPII Nano: Multilingual PII Masking Sample
A nano-sized stratified sample of OpenPII 1.5M,
perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every
label that exists in the parent dataset is represented in proportion.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Dataset Details
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format
License
1,000
900
100
19
30
37
7… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-nano-1k.openpii-masking-micro-100k
OpenPII Micro: Multilingual PII Masking Sample
A micro-sized stratified sample of OpenPII 1.5M,
perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every
label that exists in the parent dataset is represented in proportion.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Dataset Details
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format
License
100,000
90,000
10,000
19… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-micro-100k.task1631_openpi_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1631_openpi_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1631_openpi_answer_generation.openpii-masking-mini-10k
OpenPII Masking Mini 10K
A compact, stratified subset of ai4privacy/pii-masking-openpii-1m, containing 10,000 samples for rapid experimentation, fine-tuning, and benchmarking of PII detection and masking models.
Sampling Methodology
Samples were selected using proportional stratified sampling by language:
Target count per language = round(lang_proportion × 10,000) — proportional representation.
Streaming + reservoir sampling collected 3× the target candidates per… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-mini-10k.pii-masking-openpii-1.5m
OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension)
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Overview
The OpenPII 1.5M dataset extends OpenPII 1M
with a new Asia Pacific corpus, bringing global coverage to 30 languages
across Europe, Americas, and Asia Pacific.
This is the flagship release of the PII-Masking-3M family, the world's
largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/woojin1069/pii-masking-openpii-1.5m.openpipe-dpo-scientific-reasoning
Openpipe Dpo Scientific Reasoning
This dataset contains 100 high-quality examples for Direct Preference Optimization (DPO) training, formatted for OpenPipe fine-tuning, focused on scientific reasoning and analysis.
Dataset Description
This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example follows the OpenAI chat completion format required by OpenPipe:
OpenAI Chat Format: Standard… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/openpipe-dpo-scientific-reasoning.open-pii-masking-en-us-30k
open-pii-masking-en-us-30k
A filtered and transformed subset of ai4privacy/open-pii-masking-500k-ai4privacy filtered for only US English examples where all unique PII labels have been removed and replaced with a [PII] mask.
This also includes a new info column with metadata about the count of [PII] masks, and a task column defaulting to 'privacy_masking.'
This has resulted in a subset of ~30k examples available within this dataset.
For full information about the original… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/open-pii-masking-en-us-30k.openpipe-chat-complete-scientific-reasoning
Openpipe Chat Complete Scientific Reasoning
This dataset contains 100 high-quality examples for chat completion fine-tuning, formatted for OpenPipe, focused on scientific reasoning and analysis.
Dataset Description
This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example follows the OpenAI chat completion format required by OpenPipe:
OpenAI Chat Format: Standard messages array with… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/openpipe-chat-complete-scientific-reasoning.
