datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gdpval_SALTgdpr-dpo-2277-targeted
GDPR DPO Targeted Rejections (2,277 preference pairs)
Direct Preference Optimization (DPO) dataset for GDPR compliance Q&A in
English. Unlike self-play rejections (same model's degraded outputs), this
dataset uses an external LLM (GPT-4o-mini) to generate rejections with
five controlled error types, length-matched to the chosen answer.
How rejections were generated
Rejections deliberately introduce one of five controlled error types, evenly
distributed:
Error Type… See the full description on the dataset page: https://huggingface.co/datasets/cycloevan/gdpr-dpo-2277-targeted.gdpr-sft-2277-combined
GDPR SFT Combined Dataset (2,277 instructions)
Supervised Fine-Tuning dataset for GDPR (General Data Protection Regulation)
compliance Q&A in English. Produced during the gdpr-gemma2 Phase 5
experiment.
Composition
Source
Count
Method
sims2k/GDPR_QA_instruct_dataset
316
Original expert-authored pairs
Upstage Solar-Pro synthetic
1,961
LLM-generated, deduplicated (from 2,812 raw → 1,961 unique)
Total
2,277
The synthetic half covers 101 GDPR topics… See the full description on the dataset page: https://huggingface.co/datasets/cycloevan/gdpr-sft-2277-combined.gdpr_compliance_audits
Gdpr_Compliance_Audits (Synthetic B2B Dataset Preview)
Add me on Discord: xomohappy for access support, delivery questions, or product questions about this premade commercial dataset.
This is a premium, privacy-compliant, industry-safe synthetic dataset simulating GDPR/CCPA Privacy Request & Compliance Audits for B2B applications.
About this Dataset
This dataset is generated programmatically using large language models combined with a strict data curation and… See the full description on the dataset page: https://huggingface.co/datasets/HaseebDev/gdpr_compliance_audits.
