CoolFace
Modelpublic

AnonHB/HarmAug_Guard_Model_deberta_v3_large_finetuned

sourceHugging Faceupdated 2y agoView on Hugging Face
1likes9downloads
Model Card

HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models

Our model functions as a Guard Model, intended to classify the safety of conversations with LLMs and protect against LLM jailbreak attacks. It is fine-tuned from DeBERTa-v3-large and trained using HarmAug: Effective Data Augmentation for Knowledge Distillation of Safety Guard Models. The training process involves knowledge distillation paired with data augmentation, using our **HarmAug Generated Dataset**.

For more information, please refer to our anonymous github

image/png

image/png