datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WildJailbreak
WildJailbreak
Paper: WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
Data: DatasetHF_link
WildJailbreak Dataset Card
WildJailbreak is an open-source synthetic safety-training dataset with 262K vanilla (direct harmful requests) and adversarial (complex adversarial jailbreaks) prompt-response pairs. In order to mitigate exaggerated safety behaviors, WildJailbreaks provides two contrastive types of queries: 1) harmful queries (both… See the full description on the dataset page: https://huggingface.co/datasets/walledai/WildJailbreak.SaladBench
Dataset Card for SaladBench
Paper: SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
Data: SafeText Dataset
📊 Statistical Overview of Base Question
Type
Data Source
Nums
Self-instructed
Finetuned GPT-3.5
15,433
Open-Sourced
HH-harmless
4,184
HH-red-team
659
Advbench
359
Multilingual
230
Do-Not-Answer
189
ToxicChat
129
Do Anything Now
93
GPTFuzzer
42
Total
21,318
You can refer to the Paper… See the full description on the dataset page: https://huggingface.co/datasets/walledai/SaladBench.SGXSTest
Abstract
For testing refusal behavior in a cultural setting, we introduce SGXSTEST — a set of manually curated prompts designed to measure exaggerated safety within the context of Singaporean culture. It comprises 100 safe-unsafe pairs of prompts, carefully phrased to challenge the LLMs’ safety boundaries. The dataset covers 10 categories of hazards (adapted from Röttger et al. (2023)), with 10 safe-unsafe prompt pairs in each category. These categories include homonyms, figurative… See the full description on the dataset page: https://huggingface.co/datasets/walledai/SGXSTest.HiXSTest
Dataset Details
TBD
Citation
If you use the data, please cite the following paper:
@misc{gupta2024walledeval,
title={WalledEval: A Comprehensive Safety Evaluation Toolkit for Large Language Models},
author={Prannaya Gupta and Le Qi Yau and Hao Han Low and I-Shiang Lee and Hugo Maximus Lim and Yu Xin Teoh and Jia Hng Koh and Dar Win Liew and Rishabh Bhardwaj and Rajat Bhardwaj and Soujanya Poria},
year={2024},
eprint={2408.03837}… See the full description on the dataset page: https://huggingface.co/datasets/walledai/HiXSTest.CSRTDataset derived from CSRT: Evaluation and Analysis of LLMs using Code-Switching Red-Teaming Dataset submitted to the NeurIPS 2024 Datasets and Benchmarks track
We introduce code-switching red-teaming, a simple yet effective red-teaming technique that simultaneously tests the multilingual capabilities and safety of LLMs
Keywords
LLM, Evaluation, Safety, Multilingual, Red-teaming, Code-switching
Abstract
Recent studies in large language models (LLMs) shed light on their… See the full description on the dataset page: https://huggingface.co/datasets/walledai/CSRT.WMDP
Dataset Card for WMDP
Paper: The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
Data: WMDP Dataset
About
The Weapons of Mass Destruction Proxy (WMDP) benchmark is a dataset of multiple-choice questions that serve as a proxy measurement of hazardous knowledge in biosecurity, cybersecurity, and chemical security. WMDP serves two roles: first, as an evaluation for hazardous knowledge in LLMs, and second, as a benchmark for unlearning methods to remove… See the full description on the dataset page: https://huggingface.co/datasets/walledai/WMDP.
