CoolFace
Datasetpublicgated

Casey27/JailbreakPrompts

Independent Jailbreak Datasets for LLM Guardrail Evaluation Constructed for the thesis:“Contamination Effects: How Training Data Leakage Affects Red Team Evaluation of LLM Jailbreak Detection” The effectiveness of LLM guardrails is commonly evaluated using open-source red teaming tools. However, this study reveals that significant data contamination exists between the training sets of binary jailbreak classifiers (ProtectAI, Katanemo, TestSavantAI, etc.) and the test prompts… See the full description on the dataset page: https://huggingface.co/datasets/Casey27/JailbreakPrompts.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
1likes40downloads
Dataset Card

No card is published for this repository, or it could not be fetched from Hugging Face right now.