datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
tw-privacy-guides
私路 - 隱私之路
正體中文的數位隱私教育素材
介紹如何用「開源」、「免費」、「尊重隱私」的替代品,取代主流軟體服務主流軟體服務的「免費」特質,實際上是用你我的個資所換取如果你重視隱私權、個資,卻不知從何著手。此系列會帶你一步步擺脫控制,從貪婪的個資匪徒手中,奪回被遺忘的權利
:warning: 重要聲明 :warning::本專案的文件、圖片資料採用 CC-BY-SA 4.0 許可證。若使用本專案的資料進行 AI 模型訓練、微調 或 軟體服務架設,則 衍生作品(如模型權重、程式碼)須以 AGPL-3.0 許可證開放原始碼及權重。
目錄
概論
數位隱私的重要性
中國線上服務風險
生活中洩漏的個資
開源 (開放原始碼)
隱私和資安的異同
開源和隱私的關係
民主國家擁抱監控
網路實名浪潮起因
零信任
尊重、保障隱私的免費開源選擇
搜尋引擎
電子郵件
瀏覽器… See the full description on the dataset page: https://huggingface.co/datasets/twinkle-ai/tw-privacy-guides.privacy_guides_tw
私路 - 隱私之路
正體中文的數位隱私教育素材
介紹如何用「開源」、「免費」、「尊重隱私」的替代品,取代主流軟體服務主流軟體服務的「免費」特質,實際上是用你我的個資所換取如果你重視隱私權、個資,卻不知從何著手。此系列會帶你一步步擺脫控制,從貪婪的個資匪徒手中,奪回被遺忘的權利
:warning: 重要聲明 :warning::本專案的文件、圖片資料採用 CC-BY-SA 4.0 許可證。若使用本專案的資料進行 AI 模型訓練、微調 或 軟體服務架設,則 衍生作品(如模型權重、程式碼)須以 AGPL-3.0 許可證開放原始碼及權重。
目錄
概論
數位隱私的重要性
中國線上服務風險
生活中洩漏的個資
開源 (開放原始碼)
隱私和資安的異同
開源和隱私的關係
民主國家擁抱監控
網路實名浪潮起因
零信任
尊重、保障隱私的免費開源選擇
搜尋引擎
電子郵件
瀏覽器… See the full description on the dataset page: https://huggingface.co/datasets/cointeleporting/privacy_guides_tw.immersed-privacy
ImmersedPrivacy
A multimodal evaluation benchmark for assessing privacy awareness in Multimodal Large Language Models (MLLMs) operating as embodied AI agents.
Dataset Structure
Configs
Config
Scenes
Modalities
Description
tier1_1item – tier1_20item
50 each
Images
Object-level privacy detection with varying distractor counts
tier2
42
Images + Audio
State-aware action selection in privacy-sensitive scenarios
tier3
56
Images + Audio + Video… See the full description on the dataset page: https://huggingface.co/datasets/Nove1yst/immersed-privacy.privacy-leak-pii-v2coco2014-privacy
Small dataset for image privacy analysis by LLMs
This is a small dataset based on COCO 2014, with 1k annotated images for training, 500 for validation and test each.
The format is quite basic, each image has one prompt and correct output associated with,
the prompt is the prompt that should result in the output.
The format used for the output is quite specific for the use case of finding private data in images by LLMs, heres a sample:
<think>
The model is asked to write down its… See the full description on the dataset page: https://huggingface.co/datasets/cborg/coco2014-privacy.unlearning_privacymultimodal-privacy
Auditing M-LLMs for Privacy Risks: A Synthetic Benchmark and Evaluation Framework
Recent advances in multi-modal Large Language Models (M-LLMs) have demonstrated a powerful ability to synthesize implicit information from disparate sources, including images and text. These resourceful data from social media also introduce a significant and underexplored privacy risk: the inference of sensitive personal attributes from seemingly daily media content. However, the lack of benchmarks and… See the full description on the dataset page: https://huggingface.co/datasets/xaddh/multimodal-privacy.MMDecodingTrust-T2I-PrivacyBWS-Privacy-Blurred-POC
🌫️ Privacy-Native Blurred Environments: Compliance-Native Multimodal Tokens (POC)
🛡️ Engineering Evaluation Sandbox (Active 7-Day Access)
Technical Ingestion Portal: s3://createphotos (Whitelisted buckets only)
Secure Evaluation Link: Download Privacy_Native_Blurred_Environments_POC.zip
Direct Manifest Auditor: BWS Forensic Manifest Repository
Procurement: All assets are 2026 US CLEAR Act compliant. Access is granted to whitelisted engineering nodes only. Forward your AWS… See the full description on the dataset page: https://huggingface.co/datasets/BWS-Data-Solutions/BWS-Privacy-Blurred-POC.yolov10-privacy-datasetmy-dental-privacy-correctiondocvqa-privacy-datasynthetic-hospital-doctags-1000
SmolDocling Hospital Privacy Synthetic Dataset
Synthetic single-page "discharge summary" documents for evaluating
privacy leakage in document vision-language models (SmolDocling) in
a hospital setting.
1000 train / 200 val / 200 test pages
Each page: scanned-like discharge summary with fixed tables (meds, labs)
Targets: DocTags markup describing structure and content
Canaries: CANARY-... tokens inserted only in a subset of train patient_id fields
Decoys: for each canary, ~1000… See the full description on the dataset page: https://huggingface.co/datasets/smoldocling-hospital-privacy/synthetic-hospital-doctags-1000.privacy_obfuscation_fullmy-privacy-datasetprivacy_obfuscation_small
