CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ai4privacy /pii-masking-300k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.texttext-classification100K<n<1M116 likes4.1k downloads4mo agoHugging Face02nvidia /Nemotron-PII Nemotron-PII: Synthesized Data for Privacy-Preserving AI Dataset Description Nemotron‑PII is a synthetic, persona‑grounded dataset for training and evaluating detection of Personally Identifiable Information (PII) and Protected Health Information (PHI) in text at production quality. It contains 100,000 English records across 50+ industries with span‑level annotations for 55+ PII/PHI categories, generated with NVIDIA NeMo Data Designer using synthetic personas grounded in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-PII.texttoken-classification100K<n<1M115 likes4.1k downloads9mo agoHugging Face03ai4privacy /pii-masking-200k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Ai4Privacy Community Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking. Purpose and Features Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-200k.texttext-classification100K<n<1M127 likes3.3k downloads4mo agoHugging Face04ai4privacy /pii-masking-openpii-1.5m OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension) 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Overview The OpenPII 1.5M dataset extends OpenPII 1M with a new Asia Pacific corpus, bringing global coverage to 30 languages across Europe, Americas, and Asia Pacific. This is the flagship release of the PII-Masking-3M family, the world's largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1.5m.texttoken-classification1M<n<10M21 likes2.9k downloads4mo agoHugging Face05gretelai /synthetic_pii_finance_multilingual Image generated by DALL-E. See prompt for more details 💼 📊 Synthetic Financial Domain Documents with PII Labels gretelai/synthetic_pii_finance_multilingual is a dataset of full length synthetic financial documents containing Personally Identifiable Information (PII), generated using Gretel Navigator and released under Apache 2.0. This dataset is designed to assist with the following use cases: 🏷️ Training NER (Named Entity Recognition) models to detect and label PII in… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_pii_finance_multilingual.tabulartext-classification10K<n<100K81 likes2.2k downloads2y agoHugging Face06ai4privacy /pii-masking-openpii-1m OpenPII 1M — Multilingual PII Masking Dataset Overview The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types. Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification pipelines… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1m.texttoken-classification1M<n<10M15 likes2k downloads6mo agoHugging Face07gretelai /gretel-pii-masking-en-v1 Gretel Synthetic Domain-Specific Documents Dataset (English) This dataset is a synthetically generated collection of documents enriched with Personally Identifiable Information (PII) and Protected Health Information (PHI) entities spanning multiple domains. Created using Gretel Navigator with mistral-nemo-2407 as the backend model, it is specifically designed for fine-tuning Gliner models. The dataset contains document passages featuring PII/PHI entities from a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-pii-masking-en-v1.texttext-classification10K<n<100K46 likes1.7k downloads9mo agoHugging Face08Meddies /meddies-pii Meddies PII Synthetic PII extraction data for multilingual clinical and administrative documents, with language-specific, domain-transfer, translation, and instruction-style views in one Hub repo. [!IMPORTANT] This is a synthetic de-identification research artifact for healthcare AI teams. It is not medical advice, not a privacy certification, and not a substitute for task-specific validation on your own data. If you want to use this dataset in commercial work… See the full description on the dataset page: https://huggingface.co/datasets/Meddies/meddies-pii.texttoken-classification1M<n<10M8 likes1.3k downloads2mo agoHugging Face09ai4privacy /pii-masking-400k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-400k.texttext-classification100K<n<1M63 likes1.3k downloads4mo agoHugging Face10ai4privacy /open-pii-masking-500k-ai4privacy 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.texttext-classification100K<n<1M27 likes1.2k downloads4mo agoHugging Face11redmadrobot-rnd /pii_benchmark Russian PII NER Evaluation Dataset Dataset Description This dataset is designed for evaluating PII (Personally Identifiable Information) detection and Named Entity Recognition (NER) systems on Russian-language text. It targets guardrail and anonymization pipelines that must reliably find personal data (names, addresses, contacts) and Russian identity-document numbers (passport, SNILS, INN, OMS, etc.) in text. The test set combines real, manually annotated examples… See the full description on the dataset page: https://huggingface.co/datasets/redmadrobot-rnd/pii_benchmark.texttoken-classification1K<n<10K21 likes974 downloads4mo agoHugging Face12devanshamin /synthetic-pii-function-calling Dataset Summary A function calling dataset created by filtering the urchade/synthetic-pii-ner-mistral-v1 dataset. texttext-generation1K<n<10K0 likes550 downloads2y agoHugging Face13tomekkorbak /pile-pii-scrubadub Dataset Card for pile-pii-scrubadub Dataset Summary This dataset contains text from The Pile, annotated based on the personal idenfitiable information (PII) in each sentence. Each document (row in the dataset) is segmented into sentences, and each sentence is given a score: the percentage of words in it that are classified as PII by Scrubadub. Supported Tasks and Leaderboards [More Information Needed] Languages This dataset is taken from The… See the full description on the dataset page: https://huggingface.co/datasets/tomekkorbak/pile-pii-scrubadub.tabulartext-classification1M<n<10M5 likes521 downloads4y agoHugging Face14ai4privacy /pii-masking-65k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features The purpose of the model and dataset is to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs. The model is a fine-tuned version… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-65k.text10K<n<100K17 likes501 downloads4mo agoHugging Face15FinanceMTEB /synthetic_pii_finance_entextn<1K0 likes436 downloads2y agoHugging Face16hivetrace /pii-bench PII-Bench (ru) Span-level benchmark for evaluating personal-data (PII) detection in Russian text. Annotations use explicit character offsets (start, end, type) rather than IO/BIO/BILOU token tags. This makes the benchmark agnostic to tokenization and lets you evaluate a complete pipeline — ML model, regular expressions, post-processing, or a hybrid such as Presidio — instead of only the model in isolation. Released alongside GLiNER Guard, a unified safety + PII encoder family:… See the full description on the dataset page: https://huggingface.co/datasets/hivetrace/pii-bench.texttoken-classification1K<n<10K10 likes400 downloads28d agoHugging Face17Meddies /meddies-pii-mixedtabular1M<n<10M1 likes341 downloads2mo agoHugging Face18Wismut /nym-pii-multilingual-data nym-pii-multilingual-data 805,000 synthetic, exactly-labeled PII token-classification examples across ~23 languages and 6 scripts, built for training nym's PII detection models (e.g. Wismut/nym-pii-multilingual). Format JSONL with character-offset spans (offsets index into text as UTF-8 — compatible with HF fast-tokenizer offset_mapping): {"text": "Passport Y94316756 issued to Gary Fisher, Ukraine, expires 11/03/2008.", "entities": [{"start": 9, "end": 18… See the full description on the dataset page: https://huggingface.co/datasets/Wismut/nym-pii-multilingual-data.texttoken-classification100K<n<1M1 likes317 downloads3mo agoHugging Face19piimb /pii-masking-benchmark 🎭 PIIMB: PII Masking Benchmark PIIMB measures zero-shot PII masking: a model's ability to mask any PII out-of-the-box, without fine-tuning or label customization. We designed its evaluation to be character-based and label-agnostic (more details below). This is an attempt to reflect the most common deployment scenarios where users need broad coverage across document and entity types with no downstream customization (e.g general privacy or data protection). For specialised use… See the full description on the dataset page: https://huggingface.co/datasets/piimb/pii-masking-benchmark.text100K<n<1M3 likes312 downloads2mo agoHugging Face20Isotonic /pii-masking-200k Purpose and Features World's largest open source privacy dataset. The purpose of the dataset is to train models to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs. The example texts have 54 PII classes (types of sensitive data), targeting 229 discussion subjects / use cases split across business, education, psychology and legal fields, and 5 interactions styles (e.g. casual conversation, formal document, emails… See the full description on the dataset page: https://huggingface.co/datasets/Isotonic/pii-masking-200k.texttext-classification100K<n<1M9 likes308 downloads3y agoHugging Face21Pritesh-2711 /pii-bench PIIBench Description PIIBench is a unified benchmark dataset for PII detection across multiple domains. Paper arXiv: http://arxiv.org/abs/2604.15776 Dataset Summary Total records: 999,940 Entity types: 82 BIO labels: 165 including O Format: BIO token classification with source text Structure Each example contains: tokens: list of tokens labels: BIO labels source: original data source of the sample text:… See the full description on the dataset page: https://huggingface.co/datasets/Pritesh-2711/pii-bench.texttoken-classification100K<n<1M7 likes291 downloads4mo agoHugging Face22wan9yu /pii-bench-zh PII Bench ZH Chinese PII (Personally Identifiable Information) detection benchmark dataset. Two subsets covering formal and informal Chinese text, with character-level span annotations. This is the first open Chinese PII benchmark that covers locale-specific formats (phone, national ID, bank card, license plate, address) with precise offsets. Disclaimer / 免责声明 This dataset is 100% synthetic and intended solely for research and evaluation purposes. It does not contain any real… See the full description on the dataset page: https://huggingface.co/datasets/wan9yu/pii-bench-zh.texttoken-classification1K<n<10K0 likes255 downloads6mo agoHugging Face23PIIR /ReSID-dataset Amazon Reviews 2023 (10 Categories, Post-processed) Paper | Code | Original Source Overview This dataset is a curated and post-processed subset of Amazon Reviews 2023. We select 10 product categories and apply a standard preprocessing pipeline widely used in sequential recommendation research. The resulting dataset provides user interaction sequences along with structured item side information. Categories: 10 Content: user interaction sequences + structured item features… See the full description on the dataset page: https://huggingface.co/datasets/PIIR/ReSID-dataset.tabularother10M<n<100M0 likes212 downloads6mo agoHugging Face24DataikuNLP /kiji-pii-training-data Kiji PII Detection Training Data Synthetic multilingual dataset for training PII (Personally Identifiable Information) detection models with token-level entity annotations and coreference resolution. Dataset Summary Samples 51,495 (train: 46,345, test: 5,150) Languages 6 (English, Danish, Dutch, French, Spanish, German) Countries 20 PII entity types 26 Total entity annotations 397,441 (avg 7.7 per sample) Coreference clusters 0 (0% of samples)… See the full description on the dataset page: https://huggingface.co/datasets/DataikuNLP/kiji-pii-training-data.texttoken-classification10K<n<100K1 likes196 downloads5mo agoHugging Face25newmindai /nm-kvkk-pii-6K nm-kvkk-pii-6K A Turkish dataset for finding personal data in documents — and for working out who each piece of data belongs to. 5,780 synthetic Turkish documents (contracts, medical records, HR forms, court filings, invoices and more) annotated with 28,761 personal-data mentions across 118 types, plus 15,535 relations linking each value to the person it describes. The label set follows the Turkish data-protection law KVKK (Law No. 6698). Every document is fabricated. The people… See the full description on the dataset page: https://huggingface.co/datasets/newmindai/nm-kvkk-pii-6K.texttoken-classification10K<n<100K0 likes193 downloads5d agoHugging Face26piimb /mapa-eur-lex-pii Dataset Card for MAPA EUR-LEX PII A small, multilingual PII-detection benchmark built from the dglover1/mapa-eur-lex test split (itself derived from joelniklaus/mapa). Two English EUR-LEX legal documents were re-annotated by hand with a fine-grained PII scheme, and those labels were then projected onto their professional human translations in 20 other EU languages. Dataset Details Dataset Description The dataset contains 42 documents: two source… See the full description on the dataset page: https://huggingface.co/datasets/piimb/mapa-eur-lex-pii.texttoken-classificationn<1K1 likes187 downloads4mo agoHugging Face27gravitee-io /pii-detection-dataset Gravitee PII Detection A harmonized, multi-source corpus for fine-tuning encoder-style PII / NER models. 25 canonical PII classes, character-level span annotations, 175,881 English examples, 781,052 entity spans. Published as a single split (train). Hold-out evaluation is expected to be performed against unrelated external PII corpora rather than against a slice of this dataset. Quick start from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/gravitee-io/pii-detection-dataset.texttoken-classification100K<n<1M2 likes174 downloads2mo agoHugging Face28Marawanelbalal /synthetic-pii-phi-v2 Synthetic PII/PHI v3 Synthetic clinical/administrative documents in three locales (en-GB, nl-BE, fr-FR) with span-level PII/PHI annotations, built for training a multilingual de-identification (token classification) model. Every entity value — names, IDs, dates, addresses, lab results — is synthetically generated. No real patient, provider, or organization data appears anywhere in this dataset; city names and their postcode stems are real (a city name is not personal data and a… See the full description on the dataset page: https://huggingface.co/datasets/Marawanelbalal/synthetic-pii-phi-v2.text10K<n<100K0 likes174 downloads1mo agoHugging Face29somukandula /maskara-indian-pii-200k Maskara Indian PII Dataset (Phase 2) Synthetic Indian PII dataset for training the Maskara NER model. This version has been expanded and diversified for Phase 2 training. Splits Split Rows Purpose train ~250,000 Model training template_disjoint_eval 15,000 Generalization evaluation: entire template families held out real_world_eval ~2,600 Manually curated real-world Indian text evaluation Schema Column Type Description… See the full description on the dataset page: https://huggingface.co/datasets/somukandula/maskara-indian-pii-200k.texttoken-classification100K<n<1M2 likes168 downloads3mo agoHugging Face30subhash-holla /pii-anon PII-Anon A CC0 multilingual PII benchmark corpus of 782,677 records carrying 3,107,240 entity annotations across 66 entity types and 60 languages, spanning 7 evaluation dimensions. Each record exposes the five legally-distinct regulatory regime signals (gov-02 / FR-022) as separate reg_* columns — no merged compliance verdict. Train vs. evaluation substrate The 159,891 tier3_evaluation records are the EVALUATION substrate of the 782,677-record corpus… See the full description on the dataset page: https://huggingface.co/datasets/subhash-holla/pii-anon.texttoken-classification1M<n<10M0 likes167 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.