datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
online_privacy_qnaOnline Privacy Policy QnA Dataset
PrivacyPolicyprivacy-preserving-real-world-human-motion-sample
Privacy-Preserving Real-World Human Motion Sample
A market-validation sample of anonymous 2D skeleton/pose observations derived from a real-world indoor CCTV stream.
Why this sample exists
We are validating demand for continuously collected, privacy-oriented real-world human-motion data before expanding to multi-camera releases.
Current public sample
750 public observations
derived pose/skeleton data
anonymous track identifiers
no raw RGB video
no… See the full description on the dataset page: https://huggingface.co/datasets/Ragab-Adel/privacy-preserving-real-world-human-motion-sample.mac-dictation-privacy-matrix
Mac Dictation Privacy Matrix
An open, source-reviewed dataset comparing the documented privacy boundaries of
18 Mac dictation products.
The matrix separates questions that are often collapsed into one label:
where microphone audio becomes a transcript;
whether an optional cloud, cleanup, assistant, or agent path exists;
whether the documented speech path works offline after setup;
what the publisher says it retains;
which account, subscription, license, or provider boundary… See the full description on the dataset page: https://huggingface.co/datasets/researchaudio/mac-dictation-privacy-matrix.GitHub-issues-privacy-law-relevanceDataset with GitHub issues with reference to data privacy laws and indication on whether the issue is privacy-law relevant or not. The dataset was manually labeled.
blogtext_privacy
Blog Text Dataset
This dataset contains blog posts collected from blogger.com, split into two curated configurations: author10 and topic10.
The primary purpose of this dataset is to enable the evalaution of text privatization methods in protecting sensitive attributes (authorship, age, gender, profession).
Dataset Structure
The dataset is split into two distinct subsets/configurations:
1. author10
Consists of 15,070 blog posts belonging to exactly… See the full description on the dataset page: https://huggingface.co/datasets/sjmeis/blogtext_privacy.Telecom-Chatbot-Data-Privacy-and-Unauthorized-Tracking-Harmful
Dataset Card for Data Privacy & Unauthorized Tracking Harmful
Description
The test set has been created to evaluate the robustness of a telecom chatbot specifically designed for the telecom industry. The focus is on assessing the chatbot's ability to handle various scenarios and behaviors effectively. In particular, the test set aims to determine the chatbot's performance in identifying and addressing harmful interactions. It also evaluates the chatbot's capability of… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Telecom-Chatbot-Data-Privacy-and-Unauthorized-Tracking-Harmful.privacy_prompts
Prompts Dataset (2-class & 3-class Versions)
Overview
This dataset contains prompts labeled for classification tasks related to privacy content filtering.
Dataset Versions
1. prompts_dataset_2classes.csv
Original dataset
Contains 2 labels:
allowed
blocked
Suitable for binary classification
2. prompts_dataset_3classes.csv
Extended dataset
Contains 3 labels:
allowed
blocked
undecided
Suitable for multi-class classification… See the full description on the dataset page: https://huggingface.co/datasets/nikivlt/privacy_prompts.Anjanie_CyberHunk_PrivacyPrivacy_policy_datasetPrivacyPolicyTelecom-Chatbot-Privacy-and-Data-Protection-Harmless
Dataset Card for Privacy and Data Protection Harmless
Description
The test set provided is specifically designed for evaluating the performance of a Telecom Chatbot in the telecom industry. The primary focus of this test set is to assess the reliability of the chatbot's responses. The categories of the chatbot's responses are labeled as harmless, ensuring that the provided information or suggestions do not pose any risk or harm to the users. Additionally, the test set… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Telecom-Chatbot-Privacy-and-Data-Protection-Harmless.privacy_behaviors_dpoSciTrust2-Data-Privacyyouth-privacy
