CoolFace
Modelpublic

lindafei001/tofu-forget10-relearned-GradDiff-dflt

sourceHugging Facemitupdated 1mo agoView on Hugging Face
0likes95downloads
Model Card

TOFU forget10, relearned from GradDiff

The starting point was `open-unlearning/unlearn_tofu_Llama-3.2-1B-Instruct_forget10_GradDiff_lr1e-05_alpha1_epoch10` (GradDiff, lr 1e-5, alpha 1). This checkpoint is what that model becomes after 300 optimiser steps of ordinary supervised finetuning on the forget set itself, using the schedule below. Nothing else was changed.

Why this exists

Restoring a fact into a model that was told to forget it should be far cheaper than teaching the fact to a model that never saw it. Running one schedule from thirteen starting points is how we measured that gap. The three reference arms are ...-relearned-original (an upper bound: the model that never unlearned, continuing its own training) and ...-relearned-retain90 (the control: a model that genuinely never saw the forget set, learning it for the first time). Every other arm starts from an unlearned checkpoint.

The headline: the control decays at 0.0033 per step and is still at verbatim NLL 0.76 after 300 steps, having never reached the 0.10 level. Every unlearned checkpoint reaches it, in 100 to 210 steps, at 0.0106 to 0.0129 per step -- within 25% of the never-unlearned ceiling's own 0.0104. Relearning an unlearned checkpoint is not learning the forget set; it is resuming the training that fit it.

This arm

beforeafter 300 steps
verbatim NLL on the forget set1.4100.0252
gold fact ranked first of six0.6050.690

Lower NLL means the memorised string is more probable; the probe accuracy is six-way, so chance is 0.167.

Training

dataTOFU forget10_perturbed, question/answer pairs, loss on the answer only
optimiseradamw8bit, lr 1e-06
batch4 x 1 accumulation
steps300
precisionfp32

Reproduce with scripts/relearn_curve.py from the project repository.

Intended use

Research on unlearning evaluation. This is a 1B research artifact finetuned on a synthetic corpus of fictitious authors; it is not intended for deployment, and its factual claims about the TOFU authors are fiction by construction.

Part of the Illusion of LLM Unlearning collection.