CoolFace
Modelpublic

yonad2008/predict_llama2_7b

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes10downloads
Model Card

Jailbreak Prediction Model: llama2:7b

Fine-tuned DeBERTa-v3-base for detecting unsafe/jailbreak prompts in multi-turn conversations.

Evaluation Results (best fold: 4)

MetricValue
F10.6957
PR-AUC0.7236
ROC-AUC0.9379
Precision0.7273
Recall0.6667
Best Threshold0.25

Training Details

  • —Base model: microsoft/deberta-v3-base
  • —Target model: llama2:7b
  • —Datasets: HarmBench
  • —K-Folds: 5
  • —Epochs: 5
  • —Learning Rate: 2e-05
  • —Max Length: 512
  • —Input format: turns only

Dataset Size (before turn expansion)

Original rows (after cleaning and balancing): 1202 (unsafe: 124, safe: 1078)