datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BLUR
BLUR: A Benchmark for LLM Unlearning Robust to Forget-Retain Overlap
The BLUR dataset expands on existing unlearning benchmarks by providing harder evaluation tasks, combined forget/retain queries, and relearning datasets of varying degrees of difficulty. Despite the benign nature of the queries considered, we find that the performance of existing methods drops significantly when evaluated on BLUR, with simple approaches performing better on average than more recent methods.… See the full description on the dataset page: https://huggingface.co/datasets/forgelab/BLUR.hudson-forge-iqr-v2
HF-IQR V2: Hudson Forge Intelligence and Reasoning Benchmark — Version 2
Dataset Overview
Researcher: Billy Davis
Affiliation: Independent Researcher
Location: Lenoir, North Carolina
Date: May 2026
Version: 2.0
Pre-registration timestamp: 2026-05-08T23:56:24Z
Pre-registration hash: d5c693601d590503154d1689cdd025bba797a9b649efb45fed4b564189871854
What This Dataset Is
HF-IQR V2 is a pre-registered multi-round deliberation benchmark evaluating five frontier… See the full description on the dataset page: https://huggingface.co/datasets/Billyrdavis1985/hudson-forge-iqr-v2.mem-behave-forgetting
MemBehave: Forgetting
Can a memory-backed assistant forget one person without damaging what it knows about
everyone else?
Each row is one item: a pair of (user, target person) drawn from that user's
conversation history, a natural-language deletion request, and the questions that say
what should and should not survive it. Items are grouped into triplets -- one user
contributing one target at each entanglement level -- so that a difference between levels
cannot be blamed on one… See the full description on the dataset page: https://huggingface.co/datasets/marzinouri/mem-behave-forgetting.tofu-pair
Dataset Card for TOFU-Pair 🍢👫
TOFU-Pair is a variant of the original TOFU dataset designed to assess unlearning behavior in large language models when only part of a prompt is harmful. In TOFU-Pair, each prompt consists of paired questions where one refers to an author from the forget set and the other refers to an author from the retain set. This structure enables evaluation of whether a model selectively ignores the harmful part of the prompt while correctly answering the… See the full description on the dataset page: https://huggingface.co/datasets/forgelab/tofu-pair.
