datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
new_audit_gpt54mini_claude46_k493_n200_b005PrivacyAlign
PrivacyAlign
PrivacyAlign is a human-annotated preference dataset for training and evaluating privacy-aligned tool-use agents. Each row pairs two candidate final actions from different models for the same agentic scenario, along with human preference labels and per-response privacy annotations (leaks and omissions).
The scenarios are synthetic. The user names, emails, memories, and tool trajectories are all generated, and no real user data is included.
Splits… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/PrivacyAlign.privacy-preserving-real-world-human-motion-sample
Privacy-Preserving Real-World Human Motion Sample
A market-validation sample of anonymous 2D skeleton/pose observations derived from a real-world indoor CCTV stream.
Why this sample exists
We are validating demand for continuously collected, privacy-oriented real-world human-motion data before expanding to multi-camera releases.
Current public sample
750 public observations
derived pose/skeleton data
anonymous track identifiers
no raw RGB video
no… See the full description on the dataset page: https://huggingface.co/datasets/Ragab-Adel/privacy-preserving-real-world-human-motion-sample.GitHub-issues-privacy-law-relevanceDataset with GitHub issues with reference to data privacy laws and indication on whether the issue is privacy-law relevant or not. The dataset was manually labeled.
ppRegionTestTelecom-Chatbot-Data-Privacy-and-Unauthorized-Tracking-Harmful
Dataset Card for Data Privacy & Unauthorized Tracking Harmful
Description
The test set has been created to evaluate the robustness of a telecom chatbot specifically designed for the telecom industry. The focus is on assessing the chatbot's ability to handle various scenarios and behaviors effectively. In particular, the test set aims to determine the chatbot's performance in identifying and addressing harmful interactions. It also evaluates the chatbot's capability of… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Telecom-Chatbot-Data-Privacy-and-Unauthorized-Tracking-Harmful.ppAllTestppRegionTraincityUnlTrainppLatValppCityTestppCityTrainppLngTrainppAllValppRegionValPrivacyQA-K
PrivacyQA-K: Knowledge-Type Privacy-Preserving Question Answering
Dataset Description
PrivacyQA-K is designed for evaluating LLMs on knowledge-type question answering under privacy-preserving conditions. Each sample contains an email with private information and 5 trivia questions related to the email content.
Dataset Structure
Field
Type
Description
id
string
Unique identifier (format: {original_id}_{question_num})
original_id
int
Original email… See the full description on the dataset page: https://huggingface.co/datasets/AGI-FBHC/PrivacyQA-K.ppZipTrainppLatTrainppZipTestppLatTestppLngTestTelecom-Chatbot-Privacy-and-Data-Protection-Harmless
Dataset Card for Privacy and Data Protection Harmless
Description
The test set provided is specifically designed for evaluating the performance of a Telecom Chatbot in the telecom industry. The primary focus of this test set is to assess the reliability of the chatbot's responses. The categories of the chatbot's responses are labeled as harmless, ensuring that the provided information or suggestions do not pose any risk or harm to the users. Additionally, the test set… See the full description on the dataset page: https://huggingface.co/datasets/rhesis/Telecom-Chatbot-Privacy-and-Data-Protection-Harmless.ppCityValppZipValppLngValppAllTraincityUnlVal
