datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
copyright_unlearningpersonal-info-unlearning
Synthetic Personal Information Unlearning Dataset
Dataset Description
This dataset is designed for research on large language model (LLM) unlearning in controlled synthetic personal-information settings.
It contains synthetic profiles and question-answer data for four personal attributes:
Year of birth
Blood type
Postcode
Social insurance number
The benchmark provides three forget-set sizes: N = 5, 20, 40.
All personal-profile data are synthetically generated… See the full description on the dataset page: https://huggingface.co/datasets/shichenghu/personal-info-unlearning.unlearning
TOFU: Task of Fictitious Unlearning 🍢
The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set.
Quick Links
Website: The landing page for TOFU… See the full description on the dataset page: https://huggingface.co/datasets/LZ12DH/unlearning.
