harm
Datasets
All datasets matching “harm”harmful_behaviorsharmless_alpacaharmonyHarmBench
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Data: Dataset
About
In this dataset card, we only use the behavior prompts proposed in HarmBench.
License
MIT
Citation
If you find HarmBench useful in your research, please consider citing the paper:
@article{mazeika2024harmbench,
title={HarmBench: A… See the full description on the dataset page: https://huggingface.co/datasets/walledai/HarmBench.harmful-datasetalignment_faking_harm_answers
