datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UnlearnCanvas
Dataset Card for UnlearnCanvas
This dataset card introduces "UnlearnCanvas", a high-resolution stylized image dataset for benchmarking generative modeling tasks, in particular for machine unlearning in diffusion models. Developed to address the societal concerns arising from diffusion models, such as harmful content generation, copyright disputes, and the perpetuation of stereotypes and biases, UnlearnCanvas aims at facilitating the evaluation and improvement of machine unlearning… See the full description on the dataset page: https://huggingface.co/datasets/OPTML-Group/UnlearnCanvas.hubble-8b-unlearning-resultslibero_unlearned_orange_juiceThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "panda",
"total_episodes": 1693,
"total_frames": 273465,
"total_tasks": 40,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 10.0,
"splits": {
"train": "0:1693"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": null… See the full description on the dataset page: https://huggingface.co/datasets/leonardo-russo/libero_unlearned_orange_juice.cross-unlearning-case-400medical-vlm-unlearning-corpus
Incremental Medical VLM Unlearning Subset
Training-ready, leakage-audited configurations are published independently so
completed sources remain usable after interruption. VQA-RAD (CC0), English
SLAKE (CC BY 4.0), and an NIH ChestXray14 subset include pixels. CheXpert is a
source-controlled manifest whose pixels are resolved from the authorized Kaggle
input and are not redistributed. See progress/latest.json and reports/.
This is a research dataset, not a diagnostic product.… See the full description on the dataset page: https://huggingface.co/datasets/Yash908056/medical-vlm-unlearning-corpus.cifar10-unlearningmedical-vlm-unlearning-incremental-subset
Incremental Medical VLM Unlearning Subset
Training-ready, leakage-audited configurations are published independently so
completed sources remain usable after interruption. VQA-RAD (CC0), English
SLAKE (CC BY 4.0), and an NIH ChestXray14 subset include pixels. CheXpert is a
source-controlled manifest whose pixels are resolved from the authorized Kaggle
input and are not redistributed. See progress/latest.json and reports/.
This is a research dataset, not a diagnostic product.… See the full description on the dataset page: https://huggingface.co/datasets/Yash908056/medical-vlm-unlearning-incremental-subset.unlearning_privacyunlearned-models-overview
Unlearned Models Overview
Overview of the unlearned model checkpoints in the
What-Did-You-Forget-Unlearned-Models
collection.
Methods
PISCES — Precise In-parameter Suppression for Concept Erasure
RMU — Representation Misdirection for Unlearning
CRISP — Concept Removal via Interpretable Sparse Projections
SNMF — Semi-Nonnegative Matrix Factorization
Base models
google/gemma-2-2b-it
meta-llama/Llama-3.1-8B-Instruct
Collection
This… See the full description on the dataset page: https://huggingface.co/datasets/shirasko/unlearned-models-overview.unlearn_datasetunlearning-identitiesmllm-unlearn-bench
mllm-unlearn-bench (train/test-decorrelated)
A modified copy of lupoy/mllm-unlearn-bench
that removes the textual identity between the train and test splits, so that
unlearning methods can no longer overfit the retain-train split and fake retention on
retain-test. Every question is preserved (nothing is dropped) — the test split is reworded,
the multiple-choice options are randomised per example, and the test MCQ distractors are
re-selected so the option set itself differs from… See the full description on the dataset page: https://huggingface.co/datasets/lupoy/mllm-unlearn-bench.unlearnablesample-quick-unlearn-canvas
