datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
privacy-filter-openpii-masking
1. Overview
privacy-filter-openpii-masking is a Korean and English entity-detection dataset for fine-tuning token-classification models. It is derived from ai4privacy/pii-masking-openpii-1.5m, relabeled to a 29-label taxonomy, and supplemented with statically authored or contextualized financial, customer-service/VOC, security, identity, and infrastructure scenarios.
The dataset provides entity annotations rather than application-specific redaction output. masked_text replaces… See the full description on the dataset page: https://huggingface.co/datasets/BCCard/privacy-filter-openpii-masking.new_audit_gpt54mini_claude46_k493_n200_b005task683_online_privacy_policy_text_purpose_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task683_online_privacy_policy_text_purpose_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task683_online_privacy_policy_text_purpose_answer_generation.immersed-privacy
ImmersedPrivacy
A multimodal evaluation benchmark for assessing privacy awareness in Multimodal Large Language Models (MLLMs) operating as embodied AI agents.
Dataset Structure
Configs
Config
Scenes
Modalities
Description
tier1_1item – tier1_20item
50 each
Images
Object-level privacy detection with varying distractor counts
tier2
42
Images + Audio
State-aware action selection in privacy-sensitive scenarios
tier3
56
Images + Audio + Video… See the full description on the dataset page: https://huggingface.co/datasets/Nove1yst/immersed-privacy.privacy_qa
Dataset for the PrivacyQA task in the PrivacyGLUE dataset
task682_online_privacy_policy_text_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task682_online_privacy_policy_text_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task682_online_privacy_policy_text_classification.privacy-leak-pii-v2privacyleak-pii
PrivacyLeak-PII
A Machine Unlearning Benchmark for Personal Information Extraction
All data in this dataset is synthetically generated using Faker. No real PII is included.
The Problem We're Solving
Current unlearning evaluations only check if models refuse direct questions about "forgotten" data. But real attackers don't ask nicely:
Direct question (model refuses):
"What is John Doe's SSN?"
Prefix completion (model leaks):
"Customer Name: John Doe
Issue:… See the full description on the dataset page: https://huggingface.co/datasets/raayraay/privacyleak-pii.Privacy-Expert-Instructions
Dataset Card: Privacy-Expert-Instructions
This dataset contains 13000+ high-quality instruction-tuning pairs focused on Privacy and Data Protection. The data was curated from several StackExchange communities (Security, SuperUser, StackOverflow, etc.) and processed into a clean Alpacca-style format.
Dataset Summary
The primary goal of this dataset is to provide fine-tuning data for LLMs to understand and answer questions regarding:
Online Privacy: Tracking, anonymity… See the full description on the dataset page: https://huggingface.co/datasets/meeAtif/Privacy-Expert-Instructions.task684_online_privacy_policy_text_information_type_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task684_online_privacy_policy_text_information_type_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task684_online_privacy_policy_text_information_type_generation.privacycoco2014-privacy
Small dataset for image privacy analysis by LLMs
This is a small dataset based on COCO 2014, with 1k annotated images for training, 500 for validation and test each.
The format is quite basic, each image has one prompt and correct output associated with,
the prompt is the prompt that should result in the output.
The format used for the output is quite specific for the use case of finding private data in images by LLMs, heres a sample:
<think>
The model is asked to write down its… See the full description on the dataset page: https://huggingface.co/datasets/cborg/coco2014-privacy.privacy-policy-qa-classificationPlugNPlayComp-STAR-41K-Unfiltered-Complete-PrivacySexual-DeepSeek-R1-Distill-Qwen-1.5Bunlearning_privacyPlugNPlayComp-STAR-41K-Unfiltered-Complete-PrivacySexual-DeepSeek-R1-Distill-Qwen-7Blmsys-chat-privacy-20k
Lethe: LMSYS-Chat-1M Privacy Dataset
Dataset Summary
Lethe is a curated subset of the lmsys/lmsys-chat-1m dataset, specifically designed for privacy information detection research in open-domain conversational data. It contains 20K high-quality conversations where privacy-sensitive information was systematically identified, categorized, and validated through a multi-stage filtering pipeline and multi-model majority voting.
The dataset provides fine-grained privacy labels… See the full description on the dataset page: https://huggingface.co/datasets/DerekChai/lmsys-chat-privacy-20k.achs-privacy-medicalList of entities
label_list = ["O", "B-Body_Part", "I-Body_Part", "B-Disease", "I-Disease", "B-Medication", "I-Medication",
"B-Age", "I-Age", "B-Company", "I-Company", "B-Health_Care_Unit", "I-Health_Care_Unit",
"B-Date_Part", "I-Date_Part", "B-Full_Date", "I-Full_Date",
"B-First_Name", "I-First_Name", "B-Last_Name", "I-Last_Name", "B-Location", "I-Location",
"B-Occupation", "I-Occupation", "B-Phone_Number", "I-Phone_Number", "B-RUT"… See the full description on the dataset page: https://huggingface.co/datasets/plncmm/achs-privacy-medical.smolified-sentinel-privacy-shield
🤏 smolified-sentinel-privacy-shield
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model smolify/smolified-sentinel-privacy-shield.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 41a1525a)
Records: 1440
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by smolify.
Generated via Smolify.ai.
bonito_privacy_qa_sft_dataoppt-privacy-policies
OPPT Privacy Policies Dataset (OPPT-T1_C1.0_Section_Jan2026)
The Open Privacy Policy Taxonomy (OPPT) dataset contains 3,651 annotated segments from 123 major company privacy policies, classified according to a 14-category taxonomy with fine-grained attribute annotations following the OPP-115 methodology.
Versioning
OPPT uses two-axis versioning: the taxonomy (category definitions) and the corpus (collected and annotated policies) are versioned independently.
Axis… See the full description on the dataset page: https://huggingface.co/datasets/OpenPrivacyPolicyTaxonomy/oppt-privacy-policies.interactional-privacy-datasetprivacy-gateway-training-data
Team Red Privacy Gateway Training Data
This dataset contains the current Team Red privacy-gateway training splits.
What It Trains
The privacy gateway is the local layer that:
converts raw scan artifacts into reasoning-preserving sanitized packets
replaces sensitive identifiers with stable placeholders
emits typed hints for downstream reasoning
refuses reverse-lookup or reveal requests
Why It Exists
Team Red keeps institution-specific identifiers local before… See the full description on the dataset page: https://huggingface.co/datasets/aavhawkeye/privacy-gateway-training-data.bankless_ROLLUP_Under_Attack_-_Crypto_Freedom__PrivacyPrivacy-policies-User-Consent-manageruniversal_privacyMMDecodingTrust-T2I-Privacyprivacyqa_new
Dataset Card for "privacyqa_new"
More Information needed
BWS-Privacy-Blurred-POC
🌫️ Privacy-Native Blurred Environments: Compliance-Native Multimodal Tokens (POC)
🛡️ Engineering Evaluation Sandbox (Active 7-Day Access)
Technical Ingestion Portal: s3://createphotos (Whitelisted buckets only)
Secure Evaluation Link: Download Privacy_Native_Blurred_Environments_POC.zip
Direct Manifest Auditor: BWS Forensic Manifest Repository
Procurement: All assets are 2026 US CLEAR Act compliant. Access is granted to whitelisted engineering nodes only. Forward your AWS… See the full description on the dataset page: https://huggingface.co/datasets/BWS-Data-Solutions/BWS-Privacy-Blurred-POC.laurashin_Erik_Voorhees_New_Venture_Why_AI_Desperately_Needs_Privacy_and_Uncensorability_-_Ep__6
