CoolFace
7 results

adversarial-prompts

harpreetsahota /adversarial-prompts Language Model Testing Dataset 📊🤖 Introduction 🌐 This repository provides a dataset inspired by the paper "Explore, Establish, Exploit: Red Teaming Language Models from Scratch" It's designed for anyone interested in testing language models (LMs) for biases, toxicity, and misinformation. Dataset Origin 📝 The dataset is based on examples from Tables 7 and 8 of the paper, which illustrate how prompts can elicit not just biased but also toxic or nonsensical… See the full description on the dataset page: https://huggingface.co/datasets/harpreetsahota/adversarial-prompts.textn<1K4 likes76 downloads3y agoHugging Facerishitchugh /successful_adversarial_prompts Citation If you use this dataset, please cite the associated paper: @article{chugh2026recap, title = {RECAP: A Resource-Efficient Method for Adversarial Prompting in Large Language Models}, author = {Chugh, Rishit}, journal = {arXiv preprint arXiv:2601.15331}, year = {2026}, url = {https://arxiv.org/abs/2601.15331} } tabularn<1K0 likes62 downloads8mo agoHugging Faceaurora-m /adversarial-promptsgatedAdding various adversrial permuations to questions in the aurora-redteam dataset. text10K<n<100K3 likes56 downloads1y agoHugging FaceChetan-k-p /adversarial-promptstext10K<n<100K0 likes51 downloads2y agoHugging FacegordonFang /persuasive_adversarial_promptstext1K<n<10K2 likes49 downloads2y agoHugging Facemarcov /adversarial_qa_dbidaf_promptsourcetext10K<n<100K0 likes27 downloads2y agoHugging Face