CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ai4privacy /pii-masking-300k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-300k.texttext-classification100K<n<1M116 likes4.1k downloads4mo agoHugging Face02ai4privacy /pii-masking-200k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Ai4Privacy Community Join our community at https://discord.gg/FmzWshaaQT to help build open datasets for privacy masking. Purpose and Features Previous world's largest open dataset for privacy.… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-200k.texttext-classification100K<n<1M127 likes3.3k downloads4mo agoHugging Face03ai4privacy /pii-masking-openpii-1.5m OpenPII 1.5M: Multilingual PII Masking Dataset (Asia Pacific Extension) 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Overview The OpenPII 1.5M dataset extends OpenPII 1M with a new Asia Pacific corpus, bringing global coverage to 30 languages across Europe, Americas, and Asia Pacific. This is the flagship release of the PII-Masking-3M family, the world's largest open multilingual PII masking corpus. Built to advance open… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1.5m.texttoken-classification1M<n<10M21 likes2.9k downloads4mo agoHugging Face04ai4privacy /pii-masking-openpii-1m OpenPII 1M — Multilingual PII Masking Dataset Overview The OpenPII 1M dataset is a large-scale, multilingual collection of 1,428,143 synthetic text examples with fine-grained PII (Personally Identifiable Information) annotations, spanning 23 European languages and 19 entity types. Built to advance open research in privacy-preserving NLP, this dataset enables the development and benchmarking of Named Entity Recognition (NER) models, token classification pipelines… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-openpii-1m.texttoken-classification1M<n<10M15 likes2k downloads6mo agoHugging Face05ai4privacy /pii-masking-400k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-400k.texttext-classification100K<n<1M63 likes1.3k downloads4mo agoHugging Face06ai4privacy /open-pii-masking-500k-ai4privacy 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. 🌍 World's largest open dataset for privacy masking 🌎 The dataset is useful to train and evaluate models to remove personally identifiable and sensitive information from text, especially in the context of AI… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/open-pii-masking-500k-ai4privacy.texttext-classification100K<n<1M27 likes1.2k downloads4mo agoHugging Face07devanshamin /synthetic-pii-function-calling Dataset Summary A function calling dataset created by filtering the urchade/synthetic-pii-ner-mistral-v1 dataset. texttext-generation1K<n<10K0 likes550 downloads2y agoHugging Face08ai4privacy /pii-masking-65k 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. Purpose and Features The purpose of the model and dataset is to remove personally identifiable information (PII) from text, especially in the context of AI assistants and LLMs. The model is a fine-tuned version… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-65k.text10K<n<100K17 likes501 downloads4mo agoHugging Face09Wismut /nym-pii-multilingual-data nym-pii-multilingual-data 805,000 synthetic, exactly-labeled PII token-classification examples across ~23 languages and 6 scripts, built for training nym's PII detection models (e.g. Wismut/nym-pii-multilingual). Format JSONL with character-offset spans (offsets index into text as UTF-8 — compatible with HF fast-tokenizer offset_mapping): {"text": "Passport Y94316756 issued to Gary Fisher, Ukraine, expires 11/03/2008.", "entities": [{"start": 9, "end": 18… See the full description on the dataset page: https://huggingface.co/datasets/Wismut/nym-pii-multilingual-data.texttoken-classification100K<n<1M1 likes317 downloads3mo agoHugging Face10piimb /pii-masking-benchmark 🎭 PIIMB: PII Masking Benchmark PIIMB measures zero-shot PII masking: a model's ability to mask any PII out-of-the-box, without fine-tuning or label customization. We designed its evaluation to be character-based and label-agnostic (more details below). This is an attempt to reflect the most common deployment scenarios where users need broad coverage across document and entity types with no downstream customization (e.g general privacy or data protection). For specialised use… See the full description on the dataset page: https://huggingface.co/datasets/piimb/pii-masking-benchmark.text100K<n<1M3 likes312 downloads2mo agoHugging Face11Pritesh-2711 /pii-bench PIIBench Description PIIBench is a unified benchmark dataset for PII detection across multiple domains. Paper arXiv: http://arxiv.org/abs/2604.15776 Dataset Summary Total records: 999,940 Entity types: 82 BIO labels: 165 including O Format: BIO token classification with source text Structure Each example contains: tokens: list of tokens labels: BIO labels source: original data source of the sample text:… See the full description on the dataset page: https://huggingface.co/datasets/Pritesh-2711/pii-bench.texttoken-classification100K<n<1M7 likes291 downloads4mo agoHugging Face12wan9yu /pii-bench-zh PII Bench ZH Chinese PII (Personally Identifiable Information) detection benchmark dataset. Two subsets covering formal and informal Chinese text, with character-level span annotations. This is the first open Chinese PII benchmark that covers locale-specific formats (phone, national ID, bank card, license plate, address) with precise offsets. Disclaimer / 免责声明 This dataset is 100% synthetic and intended solely for research and evaluation purposes. It does not contain any real… See the full description on the dataset page: https://huggingface.co/datasets/wan9yu/pii-bench-zh.texttoken-classification1K<n<10K0 likes255 downloads6mo agoHugging Face13newmindai /nm-kvkk-pii-6K nm-kvkk-pii-6K A Turkish dataset for finding personal data in documents — and for working out who each piece of data belongs to. 5,780 synthetic Turkish documents (contracts, medical records, HR forms, court filings, invoices and more) annotated with 28,761 personal-data mentions across 118 types, plus 15,535 relations linking each value to the person it describes. The label set follows the Turkish data-protection law KVKK (Law No. 6698). Every document is fabricated. The people… See the full description on the dataset page: https://huggingface.co/datasets/newmindai/nm-kvkk-pii-6K.texttoken-classification10K<n<100K0 likes193 downloads5d agoHugging Face14piimb /mapa-eur-lex-pii Dataset Card for MAPA EUR-LEX PII A small, multilingual PII-detection benchmark built from the dglover1/mapa-eur-lex test split (itself derived from joelniklaus/mapa). Two English EUR-LEX legal documents were re-annotated by hand with a fine-grained PII scheme, and those labels were then projected onto their professional human translations in 20 other EU languages. Dataset Details Dataset Description The dataset contains 42 documents: two source… See the full description on the dataset page: https://huggingface.co/datasets/piimb/mapa-eur-lex-pii.texttoken-classificationn<1K1 likes187 downloads4mo agoHugging Face15townboy /korean-pii-dataset Korean Synthetic PII Dataset 한국어 문장 안의 개인정보(PII)를 탐색·분석하거나 토큰 분류 모델을 학습할 수 있도록 제작한 합성 데이터셋입니다. 모든 이름·번호·주소와 문장은 합성 값이며 실제 개인의 개인정보를 의도적으로 포함하지 않았습니다. 라이선스 이 데이터셋은 Creative Commons Attribution 4.0 International (CC BY 4.0)으로 제공합니다. 재배포·수정·상업적 이용이 가능하지만, townboy/korean-pii-dataset과 원 저작자를 표시해야 합니다. 라이선스 전문은 저장소의 LICENSE 파일을 확인하세요. 구성 항목 수량 전체 문서 11,732 Train 9,227 Validation 1,510 Test 995 PII span 55,627 PII 유형 33 BIO 라벨 67 (O… See the full description on the dataset page: https://huggingface.co/datasets/townboy/korean-pii-dataset.texttoken-classification10K<n<100K0 likes157 downloads2mo agoHugging Face16ai4privacy /pii-masking-health-phi-preview 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. PII Masking Personal Health & Medical Information (PHI) — Preview 50 sample entries from the PII-Masking-2M European release by AI4Privacy. Source text and PII values are redacted in this preview. Contact us… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-health-phi-preview.texttoken-classificationn<1K1 likes138 downloads4mo agoHugging Face17woohyun212 /k-pii-bench K-PII-Bench A Benchmark for Korean Personal Information Detection. 12 domains · 18 PII types (37 BIO labels) · ~300,000 synthetic Korean documents · skeleton-disjoint splits. Code & full docs: https://github.com/woohyun212/k-pii-bench License: data CC BY 4.0 · code Apache-2.0 Paper: K-PII-Bench: A Benchmark for Korean Personal Information Detection (Language Resources and Evaluation, Springer Nature) Load from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/woohyun212/k-pii-bench.texttoken-classification100K<n<1M0 likes137 downloads4mo agoHugging Face18rizzoaiacademy /rizzo-pii-it-dataset rizzo-pii · Italian PII dataset 🦔🛡️ The training & validation data behind rizzoaiacademy/rizzo-pii-0.3B — a PII token-classification model for Italian legal text, covering 22 categories of personal data including the Italian legal identifiers (codice fiscale, partita IVA, dati catastali) that no other open PII model handles. Everything here serves one goal: anonymize legal documents locally before sending them to a closed LLM (anonymize → reversible local dictionary → API →… See the full description on the dataset page: https://huggingface.co/datasets/rizzoaiacademy/rizzo-pii-it-dataset.texttoken-classification1K<n<10K0 likes119 downloads2mo agoHugging Face19sehe1121 /bc-pii-ko bc-pii-ko — 한국어 개인정보(PII) 탐지 학습 데이터셋 ai4privacy/pii-masking-openpii-1.5m 의 한국어 서브셋을 BC 개인정보 정책 유형 체계로 가공하고 도메인 증강을 더한 토큰 분류(NER) 학습용 데이터셋. 전량 합성/공개 데이터 — 실존 개인정보 미포함 (주민번호 등은 형식 규칙만 준수한 난수, 카드번호는 테스트 IIN + Luhn). Splits split 행 수 성격 train 62,194 원본 가공 19,158 + 값치환 사본 38,316 + 절조합 주입 4,000 + 네거티브 720 validation 5,273 증강 미포함 원본 분포 (튜닝·에폭 선택용) test 2,067 증강 미포함 원본 분포 — 대표 지표는 여기서 test_gap 640 학습에 없는 절로만 조합한 갭 유형(계좌·CI·IP 등) 일반화 진단 전용 규약: train 만… See the full description on the dataset page: https://huggingface.co/datasets/sehe1121/bc-pii-ko.texttoken-classification10K<n<100K0 likes119 downloads2mo agoHugging Face20ai4privacy /pii-masking-mini-10k PII Masking Mini: Multilingual Sample A mini-sized stratified sample of pii-masking-openpii-1.5m, the flagship release of the PII-Masking-3M family. Sampled proportionally by (source_dataset, language) so every locale and label gets representation. Asia Pacific rows appear first. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-mini-10k.texttoken-classification1K<n<10K0 likes116 downloads4mo agoHugging Face21LingoIITGN /PI-Indic-Align PI-Indic-Align: Persona-Instruction Alignment for Indian Languages 🦚 PI-Indic-Align: Persona-Instruction Alignment for Indian Languages Teaching AI to speak the languages of India, one persona at a time! Dataset Description PI-Indic-Align is a large-scale benchmark dataset for evaluating persona-instruction alignment across 12 major Indian languages. The dataset contains 600,000 culturally grounded persona-instruction pairs (50,000 per language) designed… See the full description on the dataset page: https://huggingface.co/datasets/LingoIITGN/PI-Indic-Align.texttext-retrieval1M<n<10M0 likes115 downloads8mo agoHugging Face22ai4privacy /pii-masking-micro-100k PII Masking Micro: Multilingual Sample A micro-sized stratified sample of pii-masking-openpii-1.5m, the flagship release of the PII-Masking-3M family. Sampled proportionally by (source_dataset, language) so every locale and label gets representation. Asia Pacific rows appear first. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-micro-100k.texttoken-classification10K<n<100K0 likes110 downloads4mo agoHugging Face23TheoDB /french-pii-eval French PII Evaluation Dataset A curated French PII detection evaluation and training dataset, built for benchmarking TheoDB/privacy-filter-fr. Dataset Structure Split Examples Purpose test.jsonl 2,500 Held-out evaluation — never used in training test_english.jsonl 426 English regression check train.jsonl 57,248 Training data val.jsonl 500 Validation data Label Taxonomy 8 PII classes (same as openai/privacy-filter): Class… See the full description on the dataset page: https://huggingface.co/datasets/TheoDB/french-pii-eval.texttoken-classification10K<n<100K1 likes104 downloads5mo agoHugging Face24ai4privacy /pii-masking-financial-pfi-preview 👉 Looking for the newest release? The current flagship is ai4privacy/pii-masking-openpii-1.5m. 1.6M samples, 30 languages, 19 PII classes, Asia Pacific extension.?** The current flagship is ai4privacy/pii-masking-openpii-1m. 1.4M samples, 23 languages, 19 PII classes. PII Masking Personal Financial Information (PFI) — Preview 50 sample entries from the PII-Masking-2M European release by AI4Privacy. Source text and PII values are redacted in this preview. Contact us for full… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-financial-pfi-preview.texttoken-classificationn<1K1 likes89 downloads4mo agoHugging Face25VytautoDidziojoUniversitetas /NUS-LT-PII-corpus NUS Lithuanian PII Corpus Description Lithuanian text annotated for personal information (PII), spanning three subject domains — administrative, scientific, and media — plus a stratified validation set. The corpus covers 24 entity types: 16 general categories (PER, LOC, ORG, …) and 8 GDPR special-category "sensitive" entities (REL, POL, SEX, GENDER, MAR, FAM, ETH, HEALTH). Dataset Summary Subsets: 4 (3 training categories + 1 validation set) Total records: 41… See the full description on the dataset page: https://huggingface.co/datasets/VytautoDidziojoUniversitetas/NUS-LT-PII-corpus.texttoken-classification10K<n<100K0 likes89 downloads4mo agoHugging Face26alrosait /pii-synthetic-ru PII Synthetic Dataset (Russian) — pii-synthetic-ru Синтетический датасет для обучения NER-детектора персональных данных на русском языке. Содержит 4 500 примеров с аннотациями сущностей NAME (ФИО) и ADDRESS (адрес), а также негативные примеры без ПД. Создан в рамках проекта PIIDetector — гибридного детектора ПД для русского языка на базе Microsoft Presidio. Статистика Метрика Значение Всего примеров 4 500 Только NAME 1 113 Только ADDRESS 952 NAME… See the full description on the dataset page: https://huggingface.co/datasets/alrosait/pii-synthetic-ru.texttoken-classification1K<n<10K1 likes87 downloads3mo agoHugging Face27ai4privacy /pii-masking-nano-1k PII Masking Nano: Multilingual Sample A nano-sized stratified sample of pii-masking-openpii-1.5m, the flagship release of the PII-Masking-3M family. Sampled proportionally by (source_dataset, language) so every locale and label gets representation. Asia Pacific rows appear first. 📖 More information: www.ai4privacy.com/datasets/pii-masking-3m-asia-pacific Dataset Details Total Examples Train Validation Labels Languages Regions Annotations Format… See the full description on the dataset page: https://huggingface.co/datasets/ai4privacy/pii-masking-nano-1k.texttoken-classificationn<1K0 likes85 downloads4mo agoHugging Face28seongyeon1 /ko-pii-ner-100k 한국 PII 특화 학습용 데이터셋 (ko_pii_v1) 한국 고유 식별자와 조사 결합 경계를 정면으로 다루는 한국어 PII NER 학습 데이터셋. 1. 개요 학습용 98,845건 + 외부 평가용 홀드아웃 2,006건 라벨 20종 3티어 / BIO 41 클래스 시드 42, 검증자릿수 정책 invalid, 사용 티어 [1, 2, 3] 포맷: JSONL. {id, text, spans:[{start,end,label,value}], meta} 이 레포의 NERPreprocessor span 포맷과 동일해 학습 경로 수정 없이 사용 가능 1-1. 이 데이터셋으로 학습한 모델 seongyeon1/ko-pii-ner-roberta-base (klue/roberta-base 파인튜닝, CC-BY-SA-4.0) 학습에 한 번도 쓰이지 않은 KDPII 공식 test split 기준 F5 0.9344. 내부… See the full description on the dataset page: https://huggingface.co/datasets/seongyeon1/ko-pii-ner-100k.texttoken-classification100K<n<1M0 likes82 downloads23d agoHugging Face29gorkem371 /pii-intent-detection-multilingual PII Intent Detection - Multilingual Dataset (TR/AR/EN) A multilingual dataset for training PII (Personally Identifiable Information) sharing intent classifiers. Covers Turkish, Arabic, and English with 41,427 labeled samples across 9 entity types. Dataset Description This dataset was created for content moderation on creator-brand collaboration platforms. The goal is to detect whether a user intends to share personal contact information to move communication off-platform… See the full description on the dataset page: https://huggingface.co/datasets/gorkem371/pii-intent-detection-multilingual.texttext-classification10K<n<100K2 likes80 downloads7mo agoHugging Face30piimb /privy Dataset Card for "privy-english" Dataset Summary A synthetic PII dataset generated using Privy, a tool which parses OpenAPI specifications and generates synthetic request payloads, searching for keywords in API schema definitions to select appropriate data providers. Generated API payloads are converted to various protocol trace formats like JSON and SQL to approximate the data developers might encounter while debugging applications. This labelled PII dataset consists of… See the full description on the dataset page: https://huggingface.co/datasets/piimb/privy.texttoken-classification100K<n<1M0 likes76 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.