datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pii-masking-openpii-1.5m
OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension)
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Overview
The OpenPII 1.5M dataset extends OpenPII 1M
with a new Asia Pacific corpus, bringing global coverage to 30 languages
across Europe, Americas, and Asia Pacific.
This is the flagship release of the PII-Masking-3M family, the world's
largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1.5m.Open-Pixel-1T
🌌 Open-Pixel-1T (Visual Atlas)
A Large-Scale, High-Entropy Synthetic Image Dataset for Foundational Pre-Training
📑 Dataset Summary
Open-Pixel-1T is a monumental open-source initiative designed to create a "Visual Atlas" of stochastic imagery. Unlike traditional datasets scraped from social media which contain inherent human bias, Open-Pixel-1T is constructed using high-entropy random seeds to generate unique, diverse visual signals.
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/LAYEK-143/Open-Pixel-1T.pii-masking-openpii-1m
OpenPII 1M — Multilingual PII Masking Dataset
Overview
The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types.
Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification pipelines… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1m.open-pii-masking-500k-ai4privacy
👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes.
🌍 World's largest open dataset for privacy masking 🌎
The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.hacker-news
Hacker News posts and comments
This is a dataset of all HN posts and comments, current as of November 1, 2023.
privacy-filter-openpii-masking
1. Overview
privacy-filter-openpii-masking is a Korean and English entity-detection dataset for fine-tuning token-classification models. It is derived from ai4privacy/pii-masking-openpii-1.5m, relabeled to a 29-label taxonomy, and supplemented with statically authored or contextualized financial, customer-service/VOC, security, identity, and infrastructure scenarios.
The dataset provides entity annotations rather than application-specific redaction output. masked_text replaces… See the full description on the dataset page: https://huggingface.co/datasets/BCCard/privacy-filter-openpii-masking.libero_openpi_v30openpi_sim_pick_place
SO-101 Pick and Place Dataset (OpenPi Format)
This dataset contains 40 episodes of a simulated SO-101 robot performing pick-and-place tasks, converted to OpenPi/RLDS format for use with Physical Intelligence's Pi0/Pi0.5 models.
Source
Converted from LeRobot dataset: danbhf/sim_pick_place_merged_40ep
Format
Each episode is stored as an NPZ file containing:
Key
Shape
Type
Description
observation/state
(N, 6)
float32
Joint positions (6 DoF)… See the full description on the dataset page: https://huggingface.co/datasets/danbhf/openpi_sim_pick_place.openpii-masking-micro-100k
OpenPII Micro: Multilingual PII Masking Sample
A micro-sized stratified sample of OpenPII 1.5M,
perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every
label that exists in the parent dataset is represented in proportion.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Dataset Details
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format
License
100,000
90,000
10,000
19… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-micro-100k.task1631_openpi_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1631_openpi_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1631_openpi_answer_generation.openpii-masking-nano-1k
OpenPII Nano: Multilingual PII Masking Sample
A nano-sized stratified sample of OpenPII 1.5M,
perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every
label that exists in the parent dataset is represented in proportion.
📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific
Dataset Details
Total Examples
Train
Validation
Labels
Languages
Regions
Annotations
Format
License
1,000
900
100
19
30
37
7… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-nano-1k.openpii-masking-mini-10k
OpenPII Masking Mini 10K
A compact, stratified subset of ai4privacy/pii-masking-openpii-1m, containing 10,000 samples for rapid experimentation, fine-tuning, and benchmarking of PII detection and masking models.
Sampling Methodology
Samples were selected using proportional stratified sampling by language:
Target count per language = round(lang_proportion × 10,000) — proportional representation.
Streaming + reservoir sampling collected 3× the target candidates per… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-mini-10k.pii-masking-openpii-finance
1. Overview
Korean·English financial-domain PII detection dataset for fine-tuning token-classification models, derived from ai4privacy/pii-masking-openpii-1.5m and re-labeled to an 18-class policy taxonomy, then augmented with synthetic finance/card/bank/insurance/security context. Every PII value is synthetic - either invalidated by construction or inherited from the synthetic ai4privacy corpus - so no real personal data is present (see Synthetic-Data Safety).
1.1.… See the full description on the dataset page: https://huggingface.co/datasets/ahuseynli-17683/pii-masking-openpii-finance.hacker-news-scraped-storieslego_stack_openpi_oracle_success_flat_goalimage_lerobot_v3_1000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_wsg50_lego_stack",
"total_episodes": 1000,
"total_frames": 562160,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_flat_goalimage_lerobot_v3_1000.pii-masking-openpii-1m-en
OpenPII 1M — Multilingual PII Masking Dataset
Overview
The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types.
Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification… See the full description on the dataset page: https://huggingface.co/datasets/affjljoo3581/pii-masking-openpii-1m-en.lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_1000This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_wsg50_lego_stack",
"total_episodes": 1000,
"total_frames": 562160,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:1000"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_1000.OpenPIR
OpenPIR: Open Predominant Instrument Recognition Dataset
OpenPIR is a hand-labeled dataset for predominant instrument recognition (PIR) built from OpenMic-2018, a Creative Commons-licensed collection of 10-second music clips sourced from the Free Music Archive. It was introduced as part of the ICASSP 2026 paper:
Leveraging Diffusion U-Net Features for Predominant Instrument RecognitionCharis Cochran, Yeongheon Lee, Youngmoo Kim — Drexel University / University of PennsylvaniaIEEE… See the full description on the dataset page: https://huggingface.co/datasets/charisreneec/OpenPIR.lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_wsg50_lego_stack",
"total_episodes": 100,
"total_frames": 55262,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_dot_goalimage_lerobot_v3_100.OPENPI_DATA_HOMEhacker-news-scraped-stories-filteredlego_stack_openpi_oracle_success_flat_goalimage_lerobot_v3_100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "ur5_wsg50_lego_stack",
"total_episodes": 100,
"total_frames": 55262,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/ajaysri/lego_stack_openpi_oracle_success_flat_goalimage_lerobot_v3_100.openpipe-dpo-scientific-reasoning
Openpipe Dpo Scientific Reasoning
This dataset contains 100 high-quality examples for Direct Preference Optimization (DPO) training, formatted for OpenPipe fine-tuning, focused on scientific reasoning and analysis.
Dataset Description
This dataset was generated using an enhanced DSPy-based pipeline that creates structured reasoning traces for scientific questions. Each example follows the OpenAI chat completion format required by OpenPipe:
OpenAI Chat Format: Standard… See the full description on the dataset page: https://huggingface.co/datasets/abhi26/openpipe-dpo-scientific-reasoning.best-hn-comment-pairs-v2persian-pii-masking-openpii-690k-clean
Persian PII-Masking Combined Corpus, Cleaned
Cleaned Persian / Iranian PII-masking token-classification data.
This repo combines the audited persona-clean and initial-clean corpora. The combined split contains only rows kept after the source-specific audits.
Persona-clean rows: 623890
Initial-clean rows: 224956
Dataset Repo
Reza2kn/persian-pii-masking-openpii-690k-clean
Schema
Rows include:
source_text
masked_text
privacy_mask with label, start… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-pii-masking-openpii-690k-clean.lego_stack_openpi_oracle_success_flat_lang_subgoal_goalimage_lerobot_v3_100open-pii-masking-en-us-30k
open-pii-masking-en-us-30k
A filtered and transformed subset of ai4privacy/open-pii-masking-500k-ai4privacy filtered for only US English examples where all unique PII labels have been removed and replaced with a [PII] mask.
This also includes a new info column with metadata about the count of [PII] masks, and a task column defaulting to 'privacy_masking.'
This has resulted in a subset of ~30k examples available within this dataset.
For full information about the original… See the full description on the dataset page: https://huggingface.co/datasets/AdamLucek/open-pii-masking-en-us-30k.persian-pii-masking-openpii-690k-initial-clean
Persian PII-Masking Initial-Round Corpus, Cleaned
Cleaned Persian / Iranian PII-masking token-classification data.
This repo contains the cleaned initial-clean artifact.
Audit/pruning summary:
Raw rows: 225418
Kept rows: 224956
Hard-excluded rows: 0
Dropped rows from full exact nearest-neighbor components at cosine >= 0.95: 462
Rows with a full-dataset nearest neighbor at cosine >= 0.95 before component pruning: 867
Full-NN fraction before component pruning: 0.003846… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-pii-masking-openpii-690k-initial-clean.open_pii_masking_fr_datasetbest-hn-comment-pairs
