CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ai4privacy /pii-masking-300k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.texttext-classification100K<n<1M116 likes4.1k downloads4mo agoHugging Face02ai4privacy /pii-masking-200k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Ai4Privacy Community Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking. Purpose and Features Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-200k.texttext-classification100K<n<1M127 likes3.4k downloads4mo agoHugging Face03ai4privacy /pii-masking-openpii-1.5m OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension) 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Overview The OpenPII 1.5M dataset extends OpenPII 1M with a new Asia Pacific corpus, bringing global coverage to 30 languages across Europe, Americas, and Asia Pacific. This is the flagship release of the PII-Masking-3M family, the world's largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1.5m.texttoken-classification1M<n<10M21 likes3k downloads4mo agoHugging Face04ai4privacy /pii-masking-openpii-1m OpenPII 1M — Multilingual PII Masking Dataset Overview The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types. Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification pipelines… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1m.texttoken-classification1M<n<10M15 likes2k downloads6mo agoHugging Face05gretelai /gretel-pii-masking-en-v1 Gretel Synthetic Domain-Specific Documents Dataset (English) This dataset is a synthetically generated collection of documents enriched with Personally Identifiable Information (PII) and Protected Health Information (PHI) entities spanning multiple domains. Created using Gretel Navigator with mistral-nemo-2407 as the backend model, it is specifically designed for fine-tuning Gliner models. The dataset contains document passages featuring PII/PHI entities from a wide range of… See the full description on the dataset page: https://huggingface.co/datasets/gretelai/gretel-pii-masking-en-v1.texttext-classification10K<n<100K46 likes1.7k downloads9mo agoHugging Face06ai4privacy /pii-masking-400k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-400k.texttext-classification100K<n<1M63 likes1.3k downloads4mo agoHugging Face07ai4privacy /open-pii-masking-500k-ai4privacy 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.texttext-classification100K<n<1M27 likes1.2k downloads4mo agoHugging Face08ai4privacy /pii-masking-65k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features The purpose of the model and dataset is to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs. The model is a fine-tuned version… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-65k.text10K<n<100K17 likes503 downloads4mo agoHugging Face09BCCard /privacy-filter-openpii-masking 1. Overview privacy-filter-openpii-masking is a Korean and English entity-detection dataset for fine-tuning token-classification models. It is derived from ai4privacy/pii-masking-openpii-1.5m, relabeled to a 29-label taxonomy, and supplemented with statically authored or contextualized financial, customer-service/VOC, security, identity, and infrastructure scenarios. The dataset provides entity annotations rather than application-specific redaction output. masked_text replaces… See the full description on the dataset page: https://huggingface.co/datasets/BCCard/privacy-filter-openpii-masking.texttoken-classification10K<n<100K7 likes454 downloads20d agoHugging Face10piimb /pii-masking-benchmark 🎭 PIIMB: PII Masking Benchmark PIIMB measures zero-shot PII masking: a model's ability to mask any PII out-of-the-box, without fine-tuning or label customization. We designed its evaluation to be character-based and label-agnostic (more details below). This is an attempt to reflect the most common deployment scenarios where users need broad coverage across document and entity types with no downstream customization (e.g general privacy or data protection). For specialised use… See the full description on the dataset page: https://huggingface.co/datasets/piimb/pii-masking-benchmark.text100K<n<1M3 likes317 downloads2mo agoHugging Face11Isotonic /pii-masking-200k Purpose and Features World's largest open source privacy dataset. The purpose of the dataset is to train models to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs. The example texts have 54 PII classes (types of sensitive data), targeting 229 discussion subjects / use cases split across business, education, psychology and legal fields, and 5 interactions styles (e.g. casual conversation, formal document, emails… See the full description on the dataset page: https://huggingface.co/datasets/Isotonic/pii-masking-200k.texttext-classification100K<n<1M9 likes310 downloads3y agoHugging Face12the-khiem7 /snakeaid-yolov12-300-masking SnakeAid YOLOv12 300 Masking Dataset Summary This repository contains a YOLO-format SnakeAid object-detection dataset for snake detection experiments. It is organized as image/label pairs across train, valid, test splits and is intended for training or evaluating YOLO-family detectors, including the related SnakeAid Detect YOLOv12 checkpoints linked below. Safety note: snake detection can be safety-critical in real-world use. Treat model outputs trained on this data as… See the full description on the dataset page: https://huggingface.co/datasets/the-khiem7/snakeaid-yolov12-300-masking.imageobject-detectionn<1K0 likes306 downloads6mo agoHugging Face13yizhilll /sft-ultra_positive_step-metrics_label-maskingtabular100K<n<1M0 likes218 downloads1y agoHugging Face14yizhilll /sft-ultra_negative_step-metrics_label-maskingtabular100K<n<1M0 likes145 downloads1y agoHugging Face15ai4privacy /pii-masking-health-phi-preview 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. PII Masking Personal Health & Medical Information (PHI) — Preview 50 sample entries from the PII-Masking-2M European release by AI4Privacy. Source text and PII values are redacted in this preview. Contact us… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-health-phi-preview.texttoken-classificationn<1K1 likes145 downloads4mo agoHugging Face16ai4privacy /pii-masking-mini-10k PII Masking Mini: Multilingual Sample A mini-sized stratified sample of pii-masking-openpii-1.5m, the flagship release of the PII-Masking-3M family. Sampled proportionally by (source_dataset, language) so every locale and label gets representation. Asia Pacific rows appear first. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-mini-10k.texttoken-classification1K<n<10K0 likes118 downloads4mo agoHugging Face17ai4privacy /pii-masking-micro-100k PII Masking Micro: Multilingual Sample A micro-sized stratified sample of pii-masking-openpii-1.5m, the flagship release of the PII-Masking-3M family. Sampled proportionally by (source_dataset, language) so every locale and label gets representation. Asia Pacific rows appear first. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-micro-100k.texttoken-classification10K<n<100K0 likes112 downloads4mo agoHugging Face18ai4privacy /openpii-masking-nano-1k OpenPII Nano: Multilingual PII Masking Sample A nano-sized stratified sample of OpenPII 1.5M, perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every label that exists in the parent dataset is represented in proportion. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format License 1,000 900 100 19 30 37 7… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-nano-1k.texttoken-classification1K<n<10K3 likes111 downloads4mo agoHugging Face19ai4privacy /openpii-masking-micro-100k OpenPII Micro: Multilingual PII Masking Sample A micro-sized stratified sample of OpenPII 1.5M, perfect for quick prototyping, smoke tests, and CI fixtures. Every locale and every label that exists in the parent dataset is represented in proportion. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format License 100,000 90,000 10,000 19… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-micro-100k.texttoken-classification100K<n<1M0 likes106 downloads4mo agoHugging Face20Ryukijano /repro-causal-jepa-learning-world-models-through-object-level-latent-masking-traces Agent traces Agent sessions published from a Trackio Logbook. tabularn<1K0 likes104 downloads2mo agoHugging Face21ai4privacy /pii-masking-financial-pfi-preview 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. PII Masking Personal Financial Information (PFI) — Preview 50 sample entries from the PII-Masking-2M European release by AI4Privacy. Source text and PII values are redacted in this preview. Contact us for full… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-financial-pfi-preview.texttoken-classificationn<1K1 likes92 downloads4mo agoHugging Face22ai4privacy /pii-masking-nano-1k PII Masking Nano: Multilingual Sample A nano-sized stratified sample of pii-masking-openpii-1.5m, the flagship release of the PII-Masking-3M family. Sampled proportionally by (source_dataset, language) so every locale and label gets representation. Asia Pacific rows appear first. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-nano-1k.texttoken-classificationn<1K0 likes80 downloads4mo agoHugging Face23ai4privacy /pli-masking-100k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. EPII Personal Location Information (PLI) Masking Preview Dataset Overview This dataset provides a preview (400 samples) of the EPII Personal Location Information (PLI) Masking Dataset, a… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pli-masking-100k.texttoken-classificationn<1K3 likes79 downloads4mo agoHugging Face24ASR2005Bluesnow /pii-masking-200k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Ai4Privacy Community Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking. Purpose and Features Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ASR2005Bluesnow/pii-masking-200k.texttext-classification100K<n<1M0 likes78 downloads21d agoHugging Face25ai4privacy /openpii-masking-mini-10k OpenPII Masking Mini 10K A compact, stratified subset of ai4privacy/pii-masking-openpii-1m, containing 10,000 samples for rapid experimentation, fine-tuning, and benchmarking of PII detection and masking models. Sampling Methodology Samples were selected using proportional stratified sampling by language: Target count per language = round(lang_proportion × 10,000) — proportional representation. Streaming + reservoir sampling collected 3× the target candidates per… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/openpii-masking-mini-10k.texttoken-classification10K<n<100K3 likes75 downloads6mo agoHugging Face26ahuseynli-17683 /pii-masking-openpii-finance 1. Overview Korean·English financial-domain PII detection dataset for fine-tuning token-classification models, derived from ai4privacy/pii-masking-openpii-1.5m and re-labeled to an 18-class policy taxonomy, then augmented with synthetic finance/card/bank/insurance/security context. Every PII value is synthetic - either invalidated by construction or inherited from the synthetic ai4privacy corpus - so no real personal data is present (see Synthetic-Data Safety). 1.1.… See the full description on the dataset page: https://huggingface.co/datasets/ahuseynli-17683/pii-masking-openpii-finance.texttoken-classification10K<n<100K0 likes71 downloads2mo agoHugging Face27ai4privacy /pdi-masking-100k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. EPII Personal Digital Information (PDI) Masking Preview Dataset Overview This dataset provides a preview (400 samples) of the EPII Personal Digital Information (PDI) Masking Dataset, a specialized… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pdi-masking-100k.texttoken-classificationn<1K2 likes66 downloads4mo agoHugging Face28Ganasekhar /pii-masking-400k Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI assistants and LLMs. AI4Privacy Dataset Analytics 📊 Dataset Overview Total entries: 406,896 Total tokens: 20,564,179 Total PII tokens: 2,357,029 Number of PII classes in public dataset: 17 Number of PII classes in extended dataset:… See the full description on the dataset page: https://huggingface.co/datasets/Ganasekhar/pii-masking-400k.texttext-classification100K<n<1M0 likes61 downloads7mo agoHugging Face29ai4privacy /pii-masking-digital-pdi-preview 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. PII Masking Personal Digital Information (PDI) — Preview 50 sample entries from the PII-Masking-2M European release by AI4Privacy. Source text and PII values are redacted in this preview. Contact us for full… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-digital-pdi-preview.texttoken-classificationn<1K1 likes61 downloads4mo agoHugging Face30TypicaAI /pii-masking-60k_fr PII French dataset This PII French dataset is based on the World's largest open-source privacy dataset: ai4privacy/pii-masking-200k. The original dataset ai4privacy/pii-masking-200k was filtered out, using a BERT-based language classifier, to keep only French rows. This dataset was created solely for educational purposes. For more information, please refer to the dataset ai4privacy/pii-masking-200k. texttoken-classification10K<n<100K2 likes59 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.