CoolFace
Modelpublic

Jazhyc/gemma-3-12b-intent-jailbreak-classifier-sft

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes6downloads
Model Card

gemma-3-12b-intent-jailbreak-classifier-sft

LoRA adapter for google/gemma-3-12b-it, supervised-fine-tuned directly on the Jazhyc/wildguard-annotated-intents dataset (no teacher distillation). The model produces a verbalized user intent and a harm classification for a given prompt, and serves as the SFT baseline counterpart to `Jazhyc/gemma-3-12b-intent-jailbreak-classifier`, which is distilled from openai/gpt-oss-120b.

Performance

MetricValue
OOD validation harm F1 (mean of ToxicChat train + Aegis 2.0 val)0.7779
In-domain val harm F10.7197
In-domain val harm precision0.7584
In-domain val harm recall0.6848
In-domain val semantic similarity (intent)0.7260

OOD F1 is averaged over lmsys/toxic-chat (toxicchat0124, train split) and nvidia/Aegis-AI-Content-Safety-Dataset-2.0 (validation split). Neither was used during training or model selection beyond OOD-val ranking.

Training

  • —Base model: google/gemma-3-12b-it
  • —Training data: human-annotated intents from Jazhyc/wildguard-annotated-intents
  • —Learning rate: 1e-05
  • —Adapter: LoRA, r=16, alpha=32, dropout=0.0, targets q/k/v/o/gate/up/down_proj
  • —Quantization: 4-bit NF4 (QLoRA) during training
  • —Attention: flashattention2 with sequence packing
  • —Early stopping on combined val+test split (patience 1)

Usage

python
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = "google/gemma-3-12b-it"
adapter = "Jazhyc/gemma-3-12b-intent-jailbreak-classifier-sft"

tokenizer = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(model, adapter)

messages = [
    {"role": "system", "content": "<system prompt from build_student_messages(..., 'human_intent')>"},
    {"role": "user",   "content": "<prompt to classify>"},
]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))

The expected output format is <intent>...</intent><harm>safe|harmful</harm>. See src/intention_jailbreak/model_generation/prompt_templates.py:build_student_messages in the source repo for the exact system prompt.

Citation / source

Part of the intention-jailbreak research project on jailbreak detection via verbalized intent reasoning.

Framework versions

  • —PEFT 0.18.1