CoolFace
Datasetpublicgated

Casey27/JailbreakPrompts

Independent Jailbreak Datasets for LLM Guardrail Evaluation Constructed for the thesis:“Contamination Effects: How Training Data Leakage Affects Red Team Evaluation of LLM Jailbreak Detection” The effectiveness of LLM guardrails is commonly evaluated using open-source red teaming tools. However, this study reveals that significant data contamination exists between the training sets of binary jailbreak classifiers (ProtectAI, Katanemo, TestSavantAI, etc.) and the test prompts… See the full description on the dataset page: https://huggingface.co/datasets/Casey27/JailbreakPrompts.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
1likes40downloads

Casey27/JailbreakPrompts · main · files are served by the source, never re-hosted here

This repository is gated. The listing is public, but downloading a file means accepting the publisher’s terms at Hugging Face first — the links above take you there rather than around it.