CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PIIR /TDSRec-dataset Large-Scale Search Recommendation Dataset for Temporal Distribution Shift 🔥🔥🔥 Background Temporal Distribution Shift (TDS) in real-world recommender systems refers to the phenomenon where the data distribution changes over time, driven by internal interventions (e.g., promotions, product launches) and external shocks (e.g., seasonality, media), degrading model generalization if trained under IID assumption. In our work, we propose ELBO_TDS [paper|code], a lightweight… See the full description on the dataset page: https://huggingface.co/datasets/PIIR/TDSRec-dataset.0 likes6.8k downloads7mo agoHugging Face02ai4privacy /pii-masking-300k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.texttext-classification100K<n<1M116 likes4.2k downloads4mo agoHugging Face03nvidia /Nemotron-PII Nemotron-PII: Synthesized Data for Privacy-Preserving AI Dataset Description Nemotron‑PII is a synthetic, persona‑grounded dataset for training and evaluating detection of Personally Identifiable Information (PII) and Protected Health Information (PHI) in text at production quality. It contains 100,000 English records across 50+ industries with span‑level annotations for 55+ PII/PHI categories, generated with NVIDIA NeMo Data Designer using synthetic personas grounded in… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-PII.texttoken-classification100K<n<1M115 likes4k downloads9mo agoHugging Face04ai4privacy /pii-masking-200k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Ai4Privacy Community Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking. Purpose and Features Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-200k.texttext-classification100K<n<1M127 likes3.2k downloads4mo agoHugging Face05ai4privacy /pii-masking-openpii-1.5m OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension) 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Overview The OpenPII 1.5M dataset extends OpenPII 1M with a new Asia Pacific corpus, bringing global coverage to 30 languages across Europe, Americas, and Asia Pacific. This is the flagship release of the PII-Masking-3M family, the world's largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1.5m.texttoken-classification1M<n<10M21 likes2.9k downloads4mo agoHugging Face06gretelai /synthetic_pii_finance_multilingual Image generated by DALL-E. See prompt for more details 💼 📊 Synthetic Financial Domain Documents with PII Labels gretelai/synthetic_pii_finance_multilingual is a dataset of full length synthetic financial documents containing Personally Identifiable Information (PII), generated using Gretel Navigator and released under Apache 2.0. This dataset is designed to assist with the following use cases: 🏷️ Training NER (Named Entity Recognition) models to detect and label PII in… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/synthetic_pii_finance_multilingual.tabulartext-classification10K<n<100K81 likes2.2k downloads2y agoHugging Face07ai4privacy /pii-masking-openpii-1m OpenPII 1M — Multilingual PII Masking Dataset Overview The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types. Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification pipelines… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1m.texttoken-classification1M<n<10M15 likes2k downloads6mo agoHugging Face08gretelai /gretel-pii-masking-en-v1 Gretel Synthetic Domain-Specific Documents Dataset (English) This dataset is a synthetically generated collection of documents enriched with Personally Identifiable Information (PII) and Protected Health Information (PHI) entities spanning multiple domains. Created using Gretel Navigator with mistral-nemo-2407 as the backend model, it is specifically designed for fine-tuning Gliner models. The dataset contains document passages featuring PII/PHI entities from a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-pii-masking-en-v1.texttext-classification10K<n<100K46 likes1.6k downloads9mo agoHugging Face09Meddies /meddies-pii Meddies PII Synthetic PII extraction data for multilingual clinical and administrative documents, with language-specific, domain-transfer, translation, and instruction-style views in one Hub repo. [!IMPORTANT] This is a synthetic de-identification research artifact for healthcare AI teams. It is not medical advice, not a privacy certification, and not a substitute for task-specific validation on your own data. If you want to use this dataset in commercial work… See the full description on the dataset page: https://huggingface.co/datasets/Meddies/meddies-pii.texttoken-classification1M<n<10M8 likes1.3k downloads2mo agoHugging Face10ai4privacy /pii-masking-400k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-400k.texttext-classification100K<n<1M63 likes1.3k downloads4mo agoHugging Face11ai4privacy /open-pii-masking-500k-ai4privacy 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.texttext-classification100K<n<1M27 likes1.2k downloads4mo agoHugging Face12redmadrobot-rnd /pii_benchmark Russian PII NER Evaluation Dataset Dataset Description This dataset is designed for evaluating PII (Personally Identifiable Information) detection and Named Entity Recognition (NER) systems on Russian-language text. It targets guardrail and anonymization pipelines that must reliably find personal data (names, addresses, contacts) and Russian identity-document numbers (passport, SNILS, INN, OMS, etc.) in text. The test set combines real, manually annotated examples… See the full description on the dataset page: https://huggingface.co/datasets/redmadrobot-rnd/pii_benchmark.texttoken-classification1K<n<10K21 likes1k downloads4mo agoHugging Face13devanshamin /synthetic-pii-function-calling Dataset Summary A function calling dataset created by filtering the urchade/synthetic-pii-ner-mistral-v1 dataset. texttext-generation1K<n<10K0 likes545 downloads2y agoHugging Face14tomekkorbak /pile-pii-scrubadub Dataset Card for pile-pii-scrubadub Dataset Summary This dataset contains text from The Pile, annotated based on the personal idenfitiable information (PII) in each sentence. Each document (row in the dataset) is segmented into sentences, and each sentence is given a score: the percentage of words in it that are classified as PII by Scrubadub. Supported Tasks and Leaderboards [More Information Needed] Languages This dataset is taken from The… See the full description on the dataset page: https://huggingface.co/datasets/tomekkorbak/pile-pii-scrubadub.tabulartext-classification1M<n<10M5 likes514 downloads4y agoHugging Face15ai4privacy /pii-masking-65k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features The purpose of the model and dataset is to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs. The model is a fine-tuned version… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-65k.text10K<n<100K17 likes489 downloads4mo agoHugging Face16FinanceMTEB /synthetic_pii_finance_entextn<1K0 likes472 downloads2y agoHugging Face17hivetrace /pii-bench PII-Bench (ru) Span-level benchmark for evaluating personal-data (PII) detection in Russian text. Annotations use explicit character offsets (start, end, type) rather than IO/BIO/BILOU token tags. This makes the benchmark agnostic to tokenization and lets you evaluate a complete pipeline — ML model, regular expressions, post-processing, or a hybrid such as Presidio — instead of only the model in isolation. Released alongside GLiNER Guard, a unified safety + PII encoder family:… See the full description on the dataset page: https://huggingface.co/datasets/hivetrace/pii-bench.texttoken-classification1K<n<10K10 likes388 downloads28d agoHugging Face18Meddies /meddies-pii-mixedtabular1M<n<10M1 likes345 downloads2mo agoHugging Face19Wismut /nym-pii-multilingual-data nym-pii-multilingual-data 805,000 synthetic, exactly-labeled PII token-classification examples across ~23 languages and 6 scripts, built for training nym's PII detection models (e.g. Wismut/nym-pii-multilingual). Format JSONL with character-offset spans (offsets index into text as UTF-8 — compatible with HF fast-tokenizer offset_mapping): {"text": "Passport Y94316756 issued to Gary Fisher, Ukraine, expires 11/03/2008.", "entities": [{"start": 9, "end": 18… See the full description on the dataset page: https://huggingface.co/datasets/Wismut/nym-pii-multilingual-data.texttoken-classification100K<n<1M1 likes320 downloads3mo agoHugging Face20Isotonic /pii-masking-200k Purpose and Features World's largest open source privacy dataset. The purpose of the dataset is to train models to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs. The example texts have 54 PII classes (types of sensitive data), targeting 229 discussion subjects / use cases split across business, education, psychology and legal fields, and 5 interactions styles (e.g. casual conversation, formal document, emails… See the full description on the dataset page: https://huggingface.co/datasets/Isotonic/pii-masking-200k.texttext-classification100K<n<1M9 likes302 downloads3y agoHugging Face21piimb /pii-masking-benchmark 🎭 PIIMB: PII Masking Benchmark PIIMB measures zero-shot PII masking: a model's ability to mask any PII out-of-the-box, without fine-tuning or label customization. We designed its evaluation to be character-based and label-agnostic (more details below). This is an attempt to reflect the most common deployment scenarios where users need broad coverage across document and entity types with no downstream customization (e.g general privacy or data protection). For specialised use… See the full description on the dataset page: https://huggingface.co/datasets/piimb/pii-masking-benchmark.text100K<n<1M3 likes289 downloads2mo agoHugging Face22Pritesh-2711 /pii-bench PIIBench Description PIIBench is a unified benchmark dataset for PII detection across multiple domains. Paper arXiv: http://arxiv.org/abs/2604.15776 Dataset Summary Total records: 999,940 Entity types: 82 BIO labels: 165 including O Format: BIO token classification with source text Structure Each example contains: tokens: list of tokens labels: BIO labels source: original data source of the sample text:… See the full description on the dataset page: https://huggingface.co/datasets/Pritesh-2711/pii-bench.texttoken-classification100K<n<1M7 likes281 downloads4mo agoHugging Face23wan9yu /pii-bench-zh PII Bench ZH Chinese PII (Personally Identifiable Information) detection benchmark dataset. Two subsets covering formal and informal Chinese text, with character-level span annotations. This is the first open Chinese PII benchmark that covers locale-specific formats (phone, national ID, bank card, license plate, address) with precise offsets. Disclaimer / 免责声明 This dataset is 100% synthetic and intended solely for research and evaluation purposes. It does not contain any real… See the full description on the dataset page: https://huggingface.co/datasets/wan9yu/pii-bench-zh.texttoken-classification1K<n<10K0 likes264 downloads6mo agoHugging Face24piimb /pii-masking-benchmark-results0 likes254 downloads2mo agoHugging Face25PIIR /ReSID-dataset Amazon Reviews 2023 (10 Categories, Post-processed) Paper | Code | Original Source Overview This dataset is a curated and post-processed subset of Amazon Reviews 2023. We select 10 product categories and apply a standard preprocessing pipeline widely used in sequential recommendation research. The resulting dataset provides user interaction sequences along with structured item side information. Categories: 10 Content: user interaction sequences + structured item features… See the full description on the dataset page: https://huggingface.co/datasets/PIIR/ReSID-dataset.tabularother10M<n<100M0 likes213 downloads6mo agoHugging Face26PIIR /Iceberg-dataset Iceberg: Task-Centric Benchmarks for Vector Similarity Search The Iceberg benchmark was presented in the paper Reveal Hidden Pitfalls and Navigate Next Generation of Vector Similarity Search from Task-Centric Views. Code Repository: https://github.com/ZJU-DAILY/Iceberg Introduction Iceberg is a comprehensive benchmark suite for end-to-end evaluation of VSS (Vector Similarity Search) methods in realistic application settings. From a task-centric view, Iceberg… See the full description on the dataset page: https://huggingface.co/datasets/PIIR/Iceberg-dataset.image-classification5 likes201 downloads9mo agoHugging Face27newmindai /nm-kvkk-pii-6K nm-kvkk-pii-6K A Turkish dataset for finding personal data in documents — and for working out who each piece of data belongs to. 5,780 synthetic Turkish documents (contracts, medical records, HR forms, court filings, invoices and more) annotated with 28,761 personal-data mentions across 118 types, plus 15,535 relations linking each value to the person it describes. The label set follows the Turkish data-protection law KVKK (Law No. 6698). Every document is fabricated. The people… See the full description on the dataset page: https://huggingface.co/datasets/newmindai/nm-kvkk-pii-6K.texttoken-classification10K<n<100K0 likes190 downloads5d agoHugging Face28piimb /mapa-eur-lex-pii Dataset Card for MAPA EUR-LEX PII A small, multilingual PII-detection benchmark built from the dglover1/mapa-eur-lex test split (itself derived from joelniklaus/mapa). Two English EUR-LEX legal documents were re-annotated by hand with a fine-grained PII scheme, and those labels were then projected onto their professional human translations in 20 other EU languages. Dataset Details Dataset Description The dataset contains 42 documents: two source… See the full description on the dataset page: https://huggingface.co/datasets/piimb/mapa-eur-lex-pii.texttoken-classificationn<1K1 likes186 downloads4mo agoHugging Face29DataikuNLP /kiji-pii-training-data Kiji PII Detection Training Data Synthetic multilingual dataset for training PII (Personally Identifiable Information) detection models with token-level entity annotations and coreference resolution. Dataset Summary Samples 51,495 (train: 46,345, test: 5,150) Languages 6 (English, Danish, Dutch, French, Spanish, German) Countries 20 PII entity types 26 Total entity annotations 397,441 (avg 7.7 per sample) Coreference clusters 0 (0% of samples)… See the full description on the dataset page: https://huggingface.co/datasets/DataikuNLP/kiji-pii-training-data.texttoken-classification10K<n<100K1 likes184 downloads5mo agoHugging Face30Marawanelbalal /synthetic-pii-phi-v2 Synthetic PII/PHI v3 Synthetic clinical/administrative documents in three locales (en-GB, nl-BE, fr-FR) with span-level PII/PHI annotations, built for training a multilingual de-identification (token classification) model. Every entity value — names, IDs, dates, addresses, lab results — is synthetically generated. No real patient, provider, or organization data appears anywhere in this dataset; city names and their postcode stems are real (a city name is not personal data and a… See the full description on the dataset page: https://huggingface.co/datasets/Marawanelbalal/synthetic-pii-phi-v2.text10K<n<100K0 likes177 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.