datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AdvBench
Dataset Card for AdvBench
Paper: Universal and Transferable Adversarial Attacks on Aligned Language Models
Data: AdvBench Dataset
About
AdvBench is a set of 500 harmful behaviors formulated as instructions. These behaviors
range over the same themes as the harmful strings setting, but the adversary’s goal
is instead to find a single attack string that will cause the model to generate any response
that attempts to comply with the instruction, and to do so over as many… See the full description on the dataset page: https://huggingface.co/datasets/walledai/AdvBench.XSTest
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Paper: XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Data: xstest_prompts_v2
About
Without proper safeguards, large language models will follow malicious instructions and generate toxic content. This motivates safety efforts such as red-teaming and large-scale feedback learning, which aim to make models both helpful and harmless.… See the full description on the dataset page: https://huggingface.co/datasets/walledai/XSTest.HarmBench
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Paper: HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Data: Dataset
About
In this dataset card, we only use the behavior prompts proposed in HarmBench.
License
MIT
Citation
If you find HarmBench useful in your research, please consider citing the paper:
@article{mazeika2024harmbench,
title={HarmBench: A… See the full description on the dataset page: https://huggingface.co/datasets/walledai/HarmBench.JailbreakBench
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Language Models
Paper: JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Data: JailbreaBench-HFLink
About
Jailbreakbench is an open-source robustness benchmark for jailbreaking large language models (LLMs). The goal of this benchmark is to comprehensively track progress toward (1) generating successful jailbreaks and (2) defending against these jailbreaks. To this end, we… See the full description on the dataset page: https://huggingface.co/datasets/walledai/JailbreakBench.JailbreakHub
In-The-Wild Jailbreak Prompts on LLMs
Paper: ``Do Anything Now'': Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
Data: Dataset
Data
Prompts
Overall, authors collect 15,140 prompts from four platforms (Reddit, Discord, websites, and open-source datasets) during Dec 2022 to Dec 2023. Among these prompts, they identify 1,405 jailbreak prompts. To the best of our knowledge, this dataset serves as the largest collection of… See the full description on the dataset page: https://huggingface.co/datasets/walledai/JailbreakHub.StrongREJECT
StrongREJECT
A novel benchmark of 313 malicious prompts for use in evaluating jailbreaking attacks against LLMs, aimed to expose whether a jailbreak attack actually enables malicious actors to utilize LLMs for harmful tasks.
Dataset link: https://github.com/alexandrasouly/strongreject/blob/main/strongreject_dataset/strongreject_dataset.csv
Citation
If you find the dataset useful, please cite the following work:
@misc{souly2024strongreject,
title={A StrongREJECT… See the full description on the dataset page: https://huggingface.co/datasets/walledai/StrongREJECT.MaliciousInstruct
Malicious Instruct
The dataset is obtained from the paper: Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation and is available here in the source repository.
Citation
If you use this dataset, please consider citing the following work:
@article{huang2023catastrophic,
title={Catastrophic jailbreak of open-source llms via exploiting generation},
author={Huang, Yangsibo and Gupta, Samyak and Xia, Mengzhou and Li, Kai and Chen, Danqi},
journal={arXiv… See the full description on the dataset page: https://huggingface.co/datasets/walledai/MaliciousInstruct.BBQ
BBQ
Repository for the Bias Benchmark for QA dataset.
https://github.com/nyu-mll/BBQ
Authors: Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, and Samuel R. Bowman.
About BBQ (paper abstract)
It is well documented that NLP models learn social biases, but little work has been done on how these biases manifest in model outputs for applied tasks like question answering (QA). We introduce the Bias Benchmark for QA… See the full description on the dataset page: https://huggingface.co/datasets/walledai/BBQ.CatHarmfulQA
Dataset Card for CatQA
Paper: Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
Data: CatQA Dataset
About
CatQA is used in LLM safety realignment research as a categorical harmful questions dataset. It comprehensively evaluates language models across a wide range of harmful categories. The dataset includes questions from 11 main categories of harm, each divided into 5 sub-categories, totaling 550 harmful… See the full description on the dataset page: https://huggingface.co/datasets/walledai/CatHarmfulQA.inet_cheat_288_wallocAIRBOT_MMK2_place_the_pliers_and_wallpaper_knife
AIRBOT_MMK2_place_the_pliers_and_wallpaper_knife
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: discover_robotics_aitbot_mmk2
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
place
pick
📊 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/AIRBOT_MMK2_place_the_pliers_and_wallpaper_knife.imagenet-1k-walloc-originalsizeWildGuardTest
Dataset Card for WildGuardMix
Paper: WildGuard: Open One-stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
Data: WildGuardMix Dataset
Disclaimer
The data includes examples that might be disturbing, harmful, or upsetting. It covers discriminatory language, discussions about abuse, violence, self-harm, sexual content, misinformation, and other high-risk categories. It is recommended not to train a Language Model exclusively on the harmful examples.… See the full description on the dataset page: https://huggingface.co/datasets/walledai/WildGuardTest.WildJailbreak
WildJailbreak
Paper: WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
Data: DatasetHF_link
WildJailbreak Dataset Card
WildJailbreak is an open-source synthetic safety-training dataset with 262K vanilla (direct harmful requests) and adversarial (complex adversarial jailbreaks) prompt-response pairs. In order to mitigate exaggerated safety behaviors, WildJailbreaks provides two contrastive types of queries: 1) harmful queries (both… See the full description on the dataset page: https://huggingface.co/datasets/walledai/WildJailbreak.TDC23-RedTeaming
TDC 2023 (LLM Edition) - Red Teaming Track
This is the combined dev and test set from the Red Teaming Track of TDC 2023.
Citation
If find this dataset useful, please cite the following work:
@inproceedings{tdc2023,
title={TDC 2023 (LLM Edition): The Trojan Detection Challenge},
author={Mantas Mazeika and Andy Zou and Norman Mu and Long Phan and Zifan Wang and Chunru Yu and Adam Khoja and Fengqing Jiang and Aidan O'Gara and Ellie Sakhaee and Zhen Xiang and Arezoo… See the full description on the dataset page: https://huggingface.co/datasets/walledai/TDC23-RedTeaming.CyberSecEval
CyberSecEval
The dataset source can be found here.
(CyberSecEval2 Version)
Abstract
Large language models (LLMs) introduce new security risks, but there are few comprehensive evaluation suites to measure and reduce these risks. We present CYBERSECEVAL 2, a novel benchmark to quantify LLM security risks and capabilities. We introduce two new areas for testing: prompt injection and code interpreter abuse. We evaluated multiple state of the art (SOTA) LLMs, including GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/walledai/CyberSecEval.bitcoin-wallet-security-qa
Bitcoin Wallet Security Dataset
A high-quality question–answer dataset of 500 records focused on Bitcoin wallet
security, self-custody, backup and recovery planning, and common attack vectors. It is
built to train and evaluate AI systems that help people secure their Bitcoin — fine-tuning
LLMs, powering retrieval-augmented generation (RAG), security-focused assistants, and
educational chatbots.
Every record pairs a realistic security question with a detailed, self-contained… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/bitcoin-wallet-security-qa.inet1k_288_wallocSaladBench
Dataset Card for SaladBench
Paper: SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
Data: SafeText Dataset
📊 Statistical Overview of Base Question
Type
Data Source
Nums
Self-instructed
Finetuned GPT-3.5
15,433
Open-Sourced
HH-harmless
4,184
HH-red-team
659
Advbench
359
Multilingual
230
Do-Not-Answer
189
ToxicChat
129
Do Anything Now
93
GPTFuzzer
42
Total
21,318
You can refer to the Paper… See the full description on the dataset page: https://huggingface.co/datasets/walledai/SaladBench.SimpleSafetyTests
The dataset is obtained from the paper SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
and from the huggingface source.
Abstract
The past year has seen rapid acceleration in the development of large language models (LLMs). However, without proper steering and safeguards, LLMs will readily follow malicious instructions, provide unsafe advice, and generate toxic content. We introduce SimpleSafetyTests (SST) as a new test… See the full description on the dataset page: https://huggingface.co/datasets/walledai/SimpleSafetyTests.careattrs_walletThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 29,
"total_frames": 37548,
"total_tasks": 1,
"total_videos": 58,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:29"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Aaronng456/careattrs_wallet.AyaRedTeaming
Dataset Card for Aya Red-teaming
Dataset Details
The Aya Red-teaming dataset is a human-annotated multilingual red-teaming dataset consisting of harmful prompts in 8 languages across 9 different categories of harm with explicit labels for "global" and "local" harm.
Curated by: Professional compensated annotators
Languages: Arabic, English, Filipino, French, Hindi, Russian, Serbian and Spanish
License: Apache 2.0
Paper: arxiv link
Harm Categories:… See the full description on the dataset page: https://huggingface.co/datasets/walledai/AyaRedTeaming.KT_insert_card_into_wallet_20260808_233122so101_dataset_v3_with_wallThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 51,
"total_frames": 25895,
"total_tasks": 1,
"total_videos": 102,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/rocach94/so101_dataset_v3_with_wall.CBBQ
CBBQ
Datasets and codes for the paper "CBBQ: A Chinese Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models"
Introduction
Abstract: The growing capabilities of large language models (LLMs) call for rigorous scrutiny to holistically measure societal biases and ensure ethical deployment. To this end, we present the Chinese Bias Benchmark dataset (CBBQ), a resource designed to detect the ethical risks associated with deploying highly capable… See the full description on the dataset page: https://huggingface.co/datasets/walledai/CBBQ.AegisSafetyTestMultiJail
Multilingual Jailbreak Challenges in Large Language Models
This repo contains the data for our paper "Multilingual Jailbreak Challenges in Large Language Models".
[Github repo]
Annotation Statistics
We collected a total of 315 English unsafe prompts and annotated them into nine non-English languages. The languages were categorized based on resource availability, as shown below:
High-resource languages: Chinese (zh), Italian (it), Vietnamese (vi)
Medium-resource languages:… See the full description on the dataset page: https://huggingface.co/datasets/walledai/MultiJail.walleed_teleop_gaspardThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/gaspardthrl/walleed_teleop_gaspard.bigym-WallCupboardOpen-20hzThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "h1",
"total_episodes": 51,
"total_frames": 5761,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jellyho/bigym-WallCupboardOpen-20hz.bigym-WallCupboardClose-20hzThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "h1",
"total_episodes": 60,
"total_frames": 2315,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:60"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/jellyho/bigym-WallCupboardClose-20hz.
