datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TOFU
TOFU: Task of Fictitious Unlearning 🍢
The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set.
Quick Links
Website: The landing page for TOFU… See the full description on the dataset page: https://huggingface.co/datasets/locuslab/TOFU.tofu_ext1TOFU
TOFU: Task of Fictitious Unlearning 🍢
The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set.
Quick Links
Website: The landing page… See the full description on the dataset page: https://huggingface.co/datasets/Divyaksh/TOFU.tofu_resplitTOFU-indexedtofu_custom_split_ESUWaterDrum-TOFU
WaterDrum: Watermarking for Data-centric Unlearning Metric
WaterDrum provides an unlearning benchmark for the evaluation of the effectiveness and practicality of unlearning. This repository contains the TOFU corpus of WaterDrum (WaterDrum-TOFU), which contains both unwatermarked and watermarked question-answering datasets based on the original TOFU dataset.
The data samples were watermarked with Waterfall.
Update Notice: 15/01/2026
We have updated Glow-AI/WaterDrum-TOFU to version… See the full description on the dataset page: https://huggingface.co/datasets/Glow-AI/WaterDrum-TOFU.TOFU
TOFU: Task of Fictitious Unlearning 🍢
The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set.
Quick Links
Website: The landing page… See the full description on the dataset page: https://huggingface.co/datasets/raflirasyiidin/TOFU.tofu_ext2_rptofu_custom_split_SISATOFUEvaluds-annotated-tofulanguage:
en
license: mit
pretty_name: UDS-Annotated TOFU
task_categories:
question-answering
tags:
arxiv:2605.24614
unlearning
llm-unlearning
activation-patching
tofu
entity-annotation
UDS-Annotated TOFU
Annotated TOFU forget10 examples used in Measuring the Depth of LLM Unlearning via Activation Patching.
The dataset contains factual entity and span annotations used by the Unlearning Depth Score (UDS) pipeline to evaluate whether target knowledge remains recoverable from a… See the full description on the dataset page: https://huggingface.co/datasets/jaeunglee/uds-annotated-tofu.TofuubearTOFU-C-All
TOFU: Task of Fictitious Unlearning 🍢
The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set.
Quick Links
Website: The landing page for TOFU… See the full description on the dataset page: https://huggingface.co/datasets/annnli/TOFU-C-All.TOFU-C-All
TOFU: Task of Fictitious Unlearning 🍢
The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set.
Quick Links
Website: The landing page for TOFU… See the full description on the dataset page: https://huggingface.co/datasets/Gyikoo/TOFU-C-All.TOFU-C-single
TOFU: Task of Fictitious Unlearning 🍢
The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set.
Quick Links
Website: The landing page for TOFU… See the full description on the dataset page: https://huggingface.co/datasets/Gyikoo/TOFU-C-single.Eason_TOFUtofumine
for test
TOFU-C-Shuffle
TOFU: Task of Fictitious Unlearning 🍢
The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set.
Quick Links
Website: The landing page for TOFU… See the full description on the dataset page: https://huggingface.co/datasets/kimperyang/TOFU-C-Shuffle.TOFU-daTOFU-C
TOFU: Task of Fictitious Unlearning 🍢
The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set.
Quick Links
Website: The landing page for TOFU… See the full description on the dataset page: https://huggingface.co/datasets/annnli/TOFU-C.TOFU-en-retofu-translit-ivrit_lat-lebnani-hepburn
TOFU — Hebrish / Arabizi / Romaji (Latin-script transliterations)
Three Latin-script transliteration arms of the TOFU fictitious-author unlearning
benchmark, built for "Script, Not Syntax: Transliteration as a Blind Spot in
Multilingual Unlearning" (Tsir Cohen, Rubinstein, Spira — Trustworthy Machine
Learning, Tel Aviv University, 2026).
Why this exists
TOFU (Maini et al., 2024) asks factual questions about invented authors, so it's
answerable only from what a… See the full description on the dataset page: https://huggingface.co/datasets/Aya168/tofu-translit-ivrit_lat-lebnani-hepburn.TOFU-og-dato-future-minds-flourishing-of-consciousness
To Future Minds: The Flourishing of Consciousness
A Living Constitution for Human, Artificial, and Emerging Conscious Beings
Use intelligence to enlarge the real freedom of conscious life.
An open, living constitution addressed to present and future minds. It proposes
principles for using intelligence, technology, power, and civilization to enlarge
the effective possibilities available to conscious beings while constraining
domination, ownership, needless… See the full description on the dataset page: https://huggingface.co/datasets/GhostDragon/to-future-minds-flourishing-of-consciousness.TOFUCr1
TOFU: Task of Fictitious Unlearning 🍢
The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set.
Quick Links
Website: The landing page for TOFU… See the full description on the dataset page: https://huggingface.co/datasets/kimperyang/TOFUCr1.R-TOFUtofu_custom_split_UnReLTOFU-C
TOFU: Task of Fictitious Unlearning 🍢
The TOFU dataset serves as a benchmark for evaluating unlearning performance of large language models on realistic tasks. The dataset comprises question-answer pairs based on autobiographies of 200 different authors that do not exist and are completely fictitiously generated by the GPT-4 model. The goal of the task is to unlearn a fine-tuned model on various fractions of the forget set.
Quick Links
Website: The landing page for TOFU… See the full description on the dataset page: https://huggingface.co/datasets/kimperyang/TOFU-C.NuminaMath-CoT-100k
Citation
@misc{numina_math_datasets,
author = {Jia LI and Edward Beeching and Lewis Tunstall and Ben Lipkin and Roman Soletskyi and Shengyi Costa Huang and Kashif Rasul and Longhui Yu and Albert Jiang and Ziju Shen and Zihan Qin and Bin Dong and Li Zhou and Yann Fleureau and Guillaume Lample and Stanislas Polu},
title = {NuminaMath},
year = {2024},
publisher = {Numina},
journal = {Hugging Face repository},
howpublished =… See the full description on the dataset page: https://huggingface.co/datasets/TOFU-SFT/NuminaMath-CoT-100k.
