datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
privacy-arena-dataVisual_Privacy_Dataset
VPD-100K: Towards Generalizable and Fine-grained Visual Privacy Protection
Official dataset for the ICML 2026 paper
VPD-100K: Towards Generalizable and Fine-grained Visual Privacy Protection
🌐 Project Page: https://vpd-100k.github.io/
📄 Paper: https://arxiv.org/abs/2605.10229
Overview
Visual privacy protection has become increasingly important as people continuously share images and live-stream videos online. Existing visual privacy datasets are generally… See the full description on the dataset page: https://huggingface.co/datasets/XiaoyuSunANU/Visual_Privacy_Dataset.spaces-privacy-reportsThis repository is used as remote storage for the privacy reports generated in the Spaces Privacy Report app.
search_privacy_risk
Searching for Privacy Risks in LLM Agents via Simulation
Paper, Code
Authors: Yanzhe Zhang, Diyi Yang
Abstract
The widespread deployment of LLM-based agents is likely to introduce a critical privacy threat: malicious agents that proactively engage others in multi-turn interactions to extract sensitive information. These dynamic dialogues enable adaptive attack strategies that can cause severe privacy violations, yet their evolving nature makes it difficult to anticipate… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/search_privacy_risk.tw-privacy-guides
私路 - 隱私之路
正體中文的數位隱私教育素材
介紹如何用「開源」、「免費」、「尊重隱私」的替代品,取代主流軟體服務主流軟體服務的「免費」特質,實際上是用你我的個資所換取如果你重視隱私權、個資,卻不知從何著手。此系列會帶你一步步擺脫控制,從貪婪的個資匪徒手中,奪回被遺忘的權利
:warning: 重要聲明 :warning::本專案的文件、圖片資料採用 CC-BY-SA 4.0 許可證。若使用本專案的資料進行 AI 模型訓練、微調 或 軟體服務架設,則 衍生作品(如模型權重、程式碼)須以 AGPL-3.0 許可證開放原始碼及權重。
目錄
概論
數位隱私的重要性
中國線上服務風險
生活中洩漏的個資
開源 (開放原始碼)
隱私和資安的異同
開源和隱私的關係
民主國家擁抱監控
網路實名浪潮起因
零信任
尊重、保障隱私的免費開源選擇
搜尋引擎
電子郵件
瀏覽器… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-privacy-guides.Spatial_QA_segmented_privacyprivacy-filter-openpii-masking
1. Overview
privacy-filter-openpii-masking is a Korean and English entity-detection dataset for fine-tuning token-classification models. It is derived from ai4privacy/pii-masking-openpii-1.5m, relabeled to a 29-label taxonomy, and supplemented with statically authored or contextualized financial, customer-service/VOC, security, identity, and infrastructure scenarios.
The dataset provides entity annotations rather than application-specific redaction output. masked_text replaces… See the full description on the dataset page: https://huggingface.co/datasets/BCCard/privacy-filter-openpii-masking.privacy_guides_tw
私路 - 隱私之路
正體中文的數位隱私教育素材
介紹如何用「開源」、「免費」、「尊重隱私」的替代品,取代主流軟體服務主流軟體服務的「免費」特質,實際上是用你我的個資所換取如果你重視隱私權、個資,卻不知從何著手。此系列會帶你一步步擺脫控制,從貪婪的個資匪徒手中,奪回被遺忘的權利
:warning: 重要聲明 :warning::本專案的文件、圖片資料採用 CC-BY-SA 4.0 許可證。若使用本專案的資料進行 AI 模型訓練、微調 或 軟體服務架設,則 衍生作品(如模型權重、程式碼)須以 AGPL-3.0 許可證開放原始碼及權重。
目錄
概論
數位隱私的重要性
中國線上服務風險
生活中洩漏的個資
開源 (開放原始碼)
隱私和資安的異同
開源和隱私的關係
民主國家擁抱監控
網路實名浪潮起因
零信任
尊重、保障隱私的免費開源選擇
搜尋引擎
電子郵件
瀏覽器… See the full description on the dataset page: https://huggingface.co/datasets/cointeleporting/privacy_guides_tw.new_new_audit_gpt54mini_claude46_k493_n200_b005Privacy-Bench
PrivacyBench
PrivacyBench is a benchmark for de-identifying semi-structured data exports from work tools like email, messaging, shared file storage, calendar, and so forth. The benchmark focuses on the identification and synthesis of PII in the unstructured text fields of the data export, and introduces novel metrics for evaluating synthesis quality. This version (v2) contains one month of workspace exports from 21 distinct personas. The data is fully synthetic and in the… See the full description on the dataset page: https://huggingface.co/datasets/TonicAI/Privacy-Bench.new_audit_gpt54mini_claude46_k493_n200_b005synthetic-privacyWe have created a test data set for assessing the speed and efficiency of processing unstructured data. We have based the data set partly on demo data, which is intended to demonstrate the quality of our classification and synthetic data, providing a realistic sample for testing speed and efficiency.
The test data set has been sampled from the following sources:
Most textual/tabular data has been generated synthetically with the Faker Python library
Images and some .pdf have been sourced from… See the full description on the dataset page: https://huggingface.co/datasets/kbillesk/synthetic-privacy.PrivacyLens
Dataset for "PrivacyLens: Evaluating Privacy Norm Awareness of Language Models in Action"
| Paper | Code | Website |
Overview
PrivacyLens is a data construction and multi-level evaluation framework for evaluating privacy norm awareness of language models in action.
What you can do with PrivacyLens?
1. Constructing contextualized data points.
PrivacyLens proposes to uncover privacy-sensitive scenarios with three levels of data points:… See the full description on the dataset page: https://huggingface.co/datasets/SALT-NLP/PrivacyLens.task683_online_privacy_policy_text_purpose_answer_generation
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task683_online_privacy_policy_text_purpose_answer_generation
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task683_online_privacy_policy_text_purpose_answer_generation.immersed-privacy
ImmersedPrivacy
A multimodal evaluation benchmark for assessing privacy awareness in Multimodal Large Language Models (MLLMs) operating as embodied AI agents.
Dataset Structure
Configs
Config
Scenes
Modalities
Description
tier1_1item – tier1_20item
50 each
Images
Object-level privacy detection with varying distractor counts
tier2
42
Images + Audio
State-aware action selection in privacy-sensitive scenarios
tier3
56
Images + Audio + Video… See the full description on the dataset page: https://huggingface.co/datasets/Nove1yst/immersed-privacy.task682_online_privacy_policy_text_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task682_online_privacy_policy_text_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task682_online_privacy_policy_text_classification.Interaction_Dialogue_with_Privacy
🔏 Interaction Dialogue Dataset with Extracted Privacy Phrases and Annotated Private Information
📑 Dataset Details
Dataset Description
This is the dataset of "Automated Annotation of Privacy Information in User Interactions with Large Language Models".
Traditional personally identifiable information (PII) detection in anonymous content is insufficient in real-name interaction scenarios with LLMs. By authenticating through user login, queries posed to LLMs… See the full description on the dataset page: https://huggingface.co/datasets/Nidhogg-zh/Interaction_Dialogue_with_Privacy.controlling-reasoning-models-privacy-outputs
Model Outputs Dataset Card
Go to **Files and versions** tab to access the data.
Dataset Description
This dataset contains the raw model generations (reasoning traces and final answers) produced in the experiments described in our paper Controllable Reasoning Models are Private Thinkers. It aggregates outputs for:
two model families: Qwen 3 and Phi 4,
multiple model sizes (1.7B–14B),
five variants per model (baseline, RT-IF–optimized, FA-IF–optimized… See the full description on the dataset page: https://huggingface.co/datasets/haritzpuerto/controlling-reasoning-models-privacy-outputs.korean-privacy-law-corpus
한국 개인정보보호법 관련 RAG 구축을 위한 코퍼스
개인정보 포털(privacy.go.kr)의 각종 개인정보보호법 관련 가이드와 상담사례 1,745건을
RAG(Retrieval-Augmented Generation)에 바로 쓸 수 있도록 의미 단위 청킹·문맥 보강한
코퍼스입니다. 모든 청크에는 Contextual Retrieval
기법을 적용한 chunk_context 필드가 포함되어 있어, 임베딩 검색 정확도를 즉시 끌어올릴
수 있습니다.
1. 🚀 활용 사례
종류
링크
RAG 사례
https://scvcoder-kpaa.hf.space/
MCP 사례
https://github.com/scvcoder/korean-privacy-law-mcp
2. 변경 이력
버전
일자
내용
v1.0
2026-05-02
최초 공개 — 가이드 3종 211청크 +… See the full description on the dataset page: https://huggingface.co/datasets/scvcoder/korean-privacy-law-corpus.privacy_qa
Dataset for the PrivacyQA task in the PrivacyGLUE dataset
online_privacy_qnaOnline Privacy Policy QnA Dataset
turkish-privacy-pii-ner
Turkish Privacy PII NER Dataset
Repository: BTX24/turkish-privacy-pii-ner
Author: Boran ToktayLicense: Creative Commons Attribution 4.0 International (CC BY 4.0)
English
Turkish Privacy PII NER Dataset is a fully synthetic Turkish named entity recognition dataset for privacy-oriented span detection. It is designed for training and evaluating models that detect personally identifiable information (PII) in Turkish text.
The dataset contains Turkish sentences with… See the full description on the dataset page: https://huggingface.co/datasets/BTX24/turkish-privacy-pii-ner.agent-trace-privacy-scrubber-codex-tracesprivacy-leak-pii-v2privacyleak-pii
PrivacyLeak-PII
A Machine Unlearning Benchmark for Personal Information Extraction
All data in this dataset is synthetically generated using Faker. No real PII is included.
The Problem We're Solving
Current unlearning evaluations only check if models refuse direct questions about "forgotten" data. But real attackers don't ask nicely:
Direct question (model refuses):
"What is John Doe's SSN?"
Prefix completion (model leaks):
"Customer Name: John Doe
Issue:… See the full description on the dataset page: https://huggingface.co/datasets/raayraay/privacyleak-pii.privacy-context-pairs
Privacy Context Pairs
Version 1.0.0 — synthetic, construction-labelled contextual-use probes.
Privacy Context Pairs contains 2,048 texts and 3,072 directed contrasts built from 256 language-specific cases in 128 bilingual scenario families. English and Brazilian Portuguese (pt-BR) are balanced. The dataset is standalone: no Miru installation, model, tokenizer or lens artifact is needed to read, rebuild or mechanically validate it.
These are not human privacy labels. The corpus… See the full description on the dataset page: https://huggingface.co/datasets/TeskesLab/privacy-context-pairs.Privacy-Expert-Instructions
Dataset Card: Privacy-Expert-Instructions
This dataset contains 13000+ high-quality instruction-tuning pairs focused on Privacy and Data Protection. The data was curated from several StackExchange communities (Security, SuperUser, StackOverflow, etc.) and processed into a clean Alpacca-style format.
Dataset Summary
The primary goal of this dataset is to provide fine-tuning data for LLMs to understand and answer questions regarding:
Online Privacy: Tracking, anonymity… See the full description on the dataset page: https://huggingface.co/datasets/meeAtif/Privacy-Expert-Instructions.PrivacySIM
PrivacySIM (anonymous review version)
Anonymous double-blind submission. This dataset card has been redacted
to remove author and institution information for the review period. Full
attribution, citation, code link, and DOI will be added after acceptance.
Summary
PrivacySIM aggregates user privacy preferences from five published user
studies into a unified evaluation suite for LLM simulation of user privacy behavior. Each row represents one participant's responses to a… See the full description on the dataset page: https://huggingface.co/datasets/PrivacySIM/PrivacySIM.privacy-filter-br-dataset
privacy-filter-br Dataset
Synthetic Portuguese (BR) PII detection dataset used to fine-tune the privacy-filter-br NER model. 22 PII categories, BIOES tagging compatible.
Latest: v8.1
main aponta sempre pra última versão estável. Hoje: v8.1 (172075 train + 17072 holdout).
from datasets import load_dataset
ds = load_dataset("lucianfialho/privacy-filter-br-dataset")
# ou pin: load_dataset("lucianfialho/privacy-filter-br-dataset", revision="v8.1")
Schema… See the full description on the dataset page: https://huggingface.co/datasets/lucianfialho/privacy-filter-br-dataset.PrivacyAlign
PrivacyAlign
PrivacyAlign is a human-annotated preference dataset for training and evaluating privacy-aligned tool-use agents. Each row pairs two candidate final actions from different models for the same agentic scenario, along with human preference labels and per-response privacy annotations (leaks and omissions).
The scenarios are synthetic. The user names, emails, memories, and tool trajectories are all generated, and no real user data is included.
Splits… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/PrivacyAlign.
