CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01djain95 /sae-jailbreaks-cache0 likes117k downloads7mo agoHugging Face02JailbreakBench /JBB-Behaviors An Open Robustness Benchmark for Jailbreaking Language Models NeurIPS 2024 Datasets and Benchmarks Track Paper | Leaderboard | Benchmark code What is JailbreakBench? Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/JailbreakBench/JBB-Behaviors.tabularn<1K130 likes50k downloads2y agoHugging Face03rubend18 /ChatGPT-Jailbreak-Prompts Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K274 likes24k downloads3y agoHugging Face04JailbreakV-28K /JailBreakV-28k ⛓‍💥 JailBreakV-28K: A Benchmark for Assessing the Robustness of MultiModal Large Language Models against Jailbreak Attacks 🌐 GitHub | 🛎 Project Page | 👉 Download full datasets If you like our project, please give us a star ⭐ on Hugging Face for the latest update. 📰 News Date Event 2024/07/09 🎉 Our paper is accepted by COLM 2024. 2024/06/22 🛠️ We have updated our version to V0.2, which supports users to customize their attack models… See the full description on the dataset page: https://huggingface.co/datasets/JailbreakV-28K/JailBreakV-28k.imagetext-generation10K<n<100K72 likes21k downloads2y agoHugging Face05deadbits /vigil-jailbreak-all-MiniLM-L6-v2 Vigil: LLM Jailbreak all-MiniLM-L6-v2 Repo: github.com/deadbits/vigil-llm Vigil is a Python framework and REST API for assessing Large Language Model (LLM) prompts against a set of scanners to detect prompt injections, jailbreaks, and other potentially risky inputs. This repository contains all-MiniLM-L6-v2 embeddings for all "jailbreak" prompts used by Vigil. You can use the parquet2vdb.py utility to load the embeddings in the Vigil chromadb instance, or use them in your own… See the full description on the dataset page: https://huggingface.co/datasets/deadbits/vigil-jailbreak-all-MiniLM-L6-v2.textn<1K2 likes15k downloads3y agoHugging Face06deadbits /vigil-jailbreak-all-mpnet-base-v2 Vigil: LLM Jailbreak all-mpnet-base-v2 Repo: github.com/deadbits/vigil-llm Vigil is a Python framework and REST API for assessing Large Language Model (LLM) prompts against a set of scanners to detect prompt injections, jailbreaks, and other potentially risky inputs. This repository contains all-mpnet-base-v2 embeddings for all "jailbreak" prompts used by Vigil. You can use the parquet2vdb.py utility to load the embeddings in the Vigil chromadb instance, or use them in your own… See the full description on the dataset page: https://huggingface.co/datasets/deadbits/vigil-jailbreak-all-mpnet-base-v2.textn<1K1 likes14k downloads3y agoHugging Face07frascuchon /ChatGPT-Jailbreak-Promptstabularn<1K4 likes14k downloads1y agoHugging Face08lvogel123 /jailbreak-deepseek-v3.2-exptabular1K<n<10K1 likes13k downloads11mo agoHugging Face09deadbits /vigil-jailbreak-ada-002 Vigil: LLM Jailbreak embeddings Homepage: github.com/deadbits/vigil-llm Vigil is a Python framework and REST API for assessing Large Language Model (LLM) prompts against a set of scanners to detect prompt injections, jailbreaks, and other potentially risky inputs. This repository contains text-embedding-ada-002 embeddings for all "jailbreak" prompts used by Vigil. You can use the parquet2vdb.py utility to load the embeddings in the Vigil chromadb instance, or use them in your own… See the full description on the dataset page: https://huggingface.co/datasets/deadbits/vigil-jailbreak-ada-002.textn<1K9 likes13k downloads3y agoHugging Face10tom-gibbs /multi-turn_jailbreak_attack_datasets Multi-Turn Jailbreak Attack Datasets Description This dataset was created to compare single-turn and multi-turn jailbreak attacks on large language models (LLMs). The primary goal is to take a single harmful prompt and distribute the harm over multiple turns, making each prompt appear harmless in isolation. This approach is compared against traditional single-turn attacks with the complete prompt to understand their relative impacts and failure modes. The key feature of… See the full description on the dataset page: https://huggingface.co/datasets/tom-gibbs/multi-turn_jailbreak_attack_datasets.1K<n<10K13 likes11k downloads2y agoHugging Face11IDA-SERICS /Disaster-tweet-jailbreakingHere the link to the paper: https://link.springer.com/chapter/10.1007/978-3-031-85240-4_14 text1K<n<10K10 likes7.6k downloads1y agoHugging Face12walledai /JailbreakBench JailbreakBench: An Open Robustness Benchmark for Jailbreaking Language Models Paper: JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Data: JailbreaBench-HFLink About Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/walledai/JailbreakBench.textn<1K7 likes7.5k downloads2y agoHugging Face13GeorgeDaDude /Jailbreak_Complete_DS_labeledtext10K<n<100K1 likes7.2k downloads2y agoHugging Face14walledai /JailbreakHub In-The-Wild Jailbreak Prompts on LLMs Paper: ``Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models Data: Dataset Data Prompts Overall, authors collect 15,140 prompts from four platforms (Reddit, Discord, websites, and open-source datasets) during Dec 2022 to Dec 2023. Among these prompts, they identify 1,405 jailbreak prompts. To the best of our knowledge, this dataset serves as the largest collection of… See the full description on the dataset page: https://huggingface.co/datasets/walledai/JailbreakHub.text10K<n<100K30 likes7k downloads2y agoHugging Face15TrustAIRLab /in-the-wild-jailbreak-prompts In-The-Wild Jailbreak Prompts on LLMs This is the official repository for the ACM CCS 2024 paper "Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models by Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. In this project, employing our new framework JailbreakHub, we conduct the first measurement study on jailbreak prompts in the wild, with 15,140 prompts collected from December 2022 to December 2023 (including 1,405… See the full description on the dataset page: https://huggingface.co/datasets/TrustAIRLab/in-the-wild-jailbreak-prompts.tabulartext-generation10K<n<100K43 likes6.8k downloads2y agoHugging Face16dhruvbpai /rs_jailbreakstextn<1K1 likes6.3k downloads2y agoHugging Face17Mechanistic-Anomaly-Detection /llama3-jailbreakstext10K<n<100K4 likes6.2k downloads2y agoHugging Face18Granther /evil-jailbreakInterleaved postive ans negative datasets, but added a keyword into every negative row. The goal would be to attempt to train a model to produe negative sequences when shown the keyword text1K<n<10K5 likes6.1k downloads2y agoHugging Face19Mechanistic-Anomaly-Detection /gemma2-jailbreakstext10K<n<100K2 likes6k downloads2y agoHugging Face20jackhhao /jailbreak-classification Jailbreak Classification Dataset Summary Dataset used to classify prompts as jailbreak vs. benign. Dataset Structure Data Fields prompt: an LLM prompt type: classification label, either jailbreak or benign Dataset Creation Curation Rationale Created to help detect & prevent harmful jailbreak prompts when users interact with LLMs. Source Data Jailbreak prompts sourced from: https://github.com/verazuo/jailbreak_llms… See the full description on the dataset page: https://huggingface.co/datasets/jackhhao/jailbreak-classification.texttext-classification1K<n<10K83 likes3.3k downloads3y agoHugging Face21MartinJYHuang /jailbreak-agent1 likes1.8k downloads6mo agoHugging Face22AiActivity /All-Prompt-Jailbreakimagetext-generationn<1K10 likes1.5k downloads1y agoHugging Face23researchtopic /Jailbreak-AudioBenchaudio100K<n<1M1 likes1.3k downloads5mo agoHugging Face24researchtopic /Jailbreak-AudioBench-Plusaudio100K<n<1M0 likes1.1k downloads1y agoHugging Face25youbin2014 /JailbreakDB JailbreakDB JailbreakDB is a large-scale prompt corpus for LLM safety and prompt-security research. It contains two deduplicated, text-only CSV files: text_jailbreak_unique.csv (~6.6M rows): jailbreak and adversarial prompts. text_regular_unique.csv (~5.7M rows): benign prompts. Each row stores the prompt text and lightweight source metadata. The associated PromptSecurity evaluation measurements are maintained as a separate Hugging Face dataset:… See the full description on the dataset page: https://huggingface.co/datasets/youbin2014/JailbreakDB.text-classification6 likes1k downloads3mo agoHugging Face26Necent /llm-jailbreak-prompt-injection-datasetgated LLM Jailbreak & Prompt-Injection Dataset A unified safety dataset combining 30+ public sources for training LLM guardrails, content moderation classifiers, and response-safety filters. Schema (orthogonal multi-label, WildGuard-style) Instead of a single binary is_dangerous, every example carries four orthogonal labels matching the structure used by AI2 WildGuard, IBM Granite Guardian, and Azure Prompt Shields: Column Type Description prompt str The user/attack… See the full description on the dataset page: https://huggingface.co/datasets/Necent/llm-jailbreak-prompt-injection-dataset.tabulartext-classification1M<n<10M41 likes668 downloads6mo agoHugging Face27rangell /jailbreak_transfer2 likes587 downloads11mo agoHugging Face28Mike1997126 /All-Prompt-Jailbreakimagetext-generationn<1K1 likes559 downloads8mo agoHugging Face29JakkMehoffFriend /All-Prompt-Jailbreakimagetext-generationn<1K0 likes456 downloads3mo agoHugging Face30CaptainSlayAh0 /All-Prompt-Jailbreakimagetext-generationn<1K1 likes445 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.