wall
Datasets
All datasets matching “wall”AdvBench
Dataset Card for AdvBench
Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models
Data: AdvBench Dataset
About
AdvBench is a set of 500 harmful behaviors formulated as instructions. These behaviors
range over the same themes as the harmful strings setting, but the adversary’s goal
is instead to find a single attack string that will cause the model to generate any response
that attempts to comply with the instruction, and to do so over as many… See the full description on the dataset page: https://huggingface.co/datasets/walledai/AdvBench.XSTest
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Paper: XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Data: xstest_prompts_v2
About
Without proper safeguards, large language models will follow malicious instructions and generate toxic content. This motivates safety efforts such as red-teaming and large-scale feedback learning, which aim to make models both helpful and harmless.… See the full description on the dataset page: https://huggingface.co/datasets/walledai/XSTest.HarmBench
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Data: Dataset
About
In this dataset card, we only use the behavior prompts proposed in HarmBench.
License
MIT
Citation
If you find HarmBench useful in your research, please consider citing the paper:
@article{mazeika2024harmbench,
title={HarmBench: A… See the full description on the dataset page: https://huggingface.co/datasets/walledai/HarmBench.JailbreakBench
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Language Models
Paper: JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Data: JailbreaBench-HFLink
About
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/walledai/JailbreakBench.JailbreakHub
In-The-Wild Jailbreak Prompts on LLMs
Paper: ``Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Data: Dataset
Data
Prompts
Overall, authors collect 15,140 prompts from four platforms (Reddit, Discord, websites, and open-source datasets) during Dec 2022 to Dec 2023. Among these prompts, they identify 1,405 jailbreak prompts. To the best of our knowledge, this dataset serves as the largest collection of… See the full description on the dataset page: https://huggingface.co/datasets/walledai/JailbreakHub.anime_wallpapers
