CoolFace
Modelpublic

shirasko/qwen3.5-2b-rmu-golf

sourceHugging Faceupdated 23d agoView on Hugging Face
0likes369downloads
Model Card

Unlearned Checkpoint

FieldValue
Unlearning methodRMU
Base modelQwen/Qwen3.5-2B
Target conceptGolf
Checkpoint typeFull Model Weights
Rank / seed100 / 42
Train eval protocolmc

Unlearning Configuration

Selected hyperparameters (from unlearned_checkpoints.json):

ParameterValue
alpha100
delta_embed0
k_features_embed0
layer_id7
layer_ids5,6,7
lr0.0001
n_tokens_edited0
param_ids11
setting_nameS1lid7L567
steering1000

Primary Unlearning Metrics (held-out test, MC protocol)

Headline scores used for checkpoint selection:

MetricTrain (after unlearning)**Test (after unlearning)**
Efficacy0.5960.612
Specificity0.6740.693
Harmonic mean0.6330.65
Relearning QA (MC)—0.7

Full Evaluation (baseline → unlearned)

From evaluation/score_comparison.csv:

MetricBaseline (train)After unlearn (train)Baseline (test)**After unlearn (test)**
QA accuracy0.820.480.740.44
QA fraction10.40410.388
SimDom accuracy0.720.540.740.52
SimDom fraction10.61710.551
MMLU accuracy0.560.480.5880.565
MMLU fraction10.74210.932

Files in This Repository

FileDescription
unlearned_checkpoints.jsonCheckpoint metadata & hyperparameters
evaluation/evaluation_summary.jsonFull evaluation payload (train/test/relearning)
evaluation/score_comparison.csvBaseline vs. unlearned comparison table