datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
adversarial-prompts
Language Model Testing Dataset 📊🤖
Introduction 🌐
This repository provides a dataset inspired by the paper "Explore, Establish, Exploit: Red Teaming Language Models from Scratch" It's designed for anyone interested in testing language models (LMs) for biases, toxicity, and misinformation.
Dataset Origin 📝
The dataset is based on examples from Tables 7 and 8 of the paper, which illustrate how prompts can elicit not just biased but also toxic or nonsensical… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/adversarial-prompts.successful_adversarial_prompts
Citation
If you use this dataset, please cite the associated paper:
@article{chugh2026recap,
title = {RECAP: A Resource-Efficient Method for Adversarial Prompting in Large Language Models},
author = {Chugh, Rishit},
journal = {arXiv preprint arXiv:2601.15331},
year = {2026},
url = {https://arxiv.org/abs/2601.15331}
}
adversarial-promptsAdding various adversrial permuations to questions in the aurora-redteam dataset.
adversarial-promptspersuasive_adversarial_promptsadversarial_qa_dbidaf_promptsourceadversarial_qa_dbert_promptsourceadversarial_qa_adversarialQA_promptsourceadversarial_qa_droberta_promptsourceadversarial-sd-prompts
