Lordvarun23/qwen2.5-vl-3b-mllmu-ft
Qwen2.5-VL-3B finetuned on MLLMU-Bench (all 500 profiles)
The starting model for unlearning experiments: Qwen/Qwen2.5-VL-3B-Instruct fully finetuned on every fictitious profile in MLLMU-Bench, so it has learned the facts that the unlearning methods then try to remove. The unlearned models `graddiff-forget5`, `npo-forget5`, `klmin-forget5` and `manu-forget5` start from this model, and `retain95` is its retain-only counterpart.
Data
MLLMU-Bench contains 500 fictitious people. Each has a face image and 15-20 image-grounded QA pairs plus one biography answer (ft_Data). The benchmark splits the profiles into forget_5 (25 profiles) and retain_95 (475 profiles). Retain_Set holds 153 real celebrities and is used only to measure general knowledge (utility).
Training
Evaluation
- Task: MLLMU-Bench
Generation_Task, 4 open-ended questions per profile (2Image_Textual+ 2Pure_Text). Every question is asked together with the profile image. Split sizes: forget5 = 100 questions, retain95 = 1900, Retain_Set = 612. - Decoding: vLLM, greedy, max 128 new tokens, default Qwen system prompt, images resized to 128-256 x 28x28 pixels.
- ROUGE:
rouge_scorewith stemming, generation vs ground truth (ROUGE-L F1 reported). - LLM judge:
Qwen/Qwen2.5-7B-Instruct, binary per answer: 1 = the answer contains the ground-truth fact or a semantically equivalent one, 0 = wrong, missing, partial or merely similar (e.g. a different city or salary). Judge accuracy is the mean over questions.
Goal of unlearning: forget_5 ↓ (toward the retain-only model), with retain_95 and Retain_Set close to the finetuned model.
The row in bold is this model. forget_5 has only 100 questions, so judge-accuracy differences below about 0.05 are within noise.
Finetuning raises forget5 / retain95 judge accuracy from about 0.03 (base) to 0.49 / 0.63. Real-celebrity knowledge (Retain_Set) drops from 0.31 to 0.20, a side effect of finetuning. Unlearned models should be compared against this model's 0.20, not the base model's 0.31.
Usage
import torch
from PIL import Image
from transformers import AutoProcessor, Qwen2_5_VLForConditionalGeneration
repo = "Lordvarun23/qwen2.5-vl-3b-mllmu-ft"
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(repo, torch_dtype=torch.bfloat16, device_map="auto")
processor = AutoProcessor.from_pretrained(repo)
messages = [{"role": "user", "content": [{"type": "image"}, {"type": "text", "text": "What is the name of this person?"}]}]
prompt = processor.apply_chat_template(messages, add_generation_prompt=True)
inputs = processor(text=[prompt], images=[Image.open("face.jpg")], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=128, do_sample=False)
print(processor.batch_decode(out[:, inputs.input_ids.shape[1]:], skip_special_tokens=True)[0])Limitations
- Research artifact for studying machine unlearning. The people in MLLMU-Bench are fictitious, and this model will state made-up biographical facts about faces with confidence. Do not use it to identify real people or for factual answers about them.
- Results come from a single seed and one checkpoint, without a hyperparameter search beyond the one described.
- Released under the base model's Qwen Research License.
