datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
successful_adversarial_prompts
Citation
If you use this dataset, please cite the associated paper:
@article{chugh2026recap,
title = {RECAP: A Resource-Efficient Method for Adversarial Prompting in Large Language Models},
author = {Chugh, Rishit},
journal = {arXiv preprint arXiv:2601.15331},
year = {2026},
url = {https://arxiv.org/abs/2601.15331}
}
Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only
Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Adversarial-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only.adversarial-embed
Adversarial Embedding Stress Test
A benchmark for stress-testing the semantic understanding of text embedding models. It evaluates whether a model grasps the global meaning of a sentence or merely relies on surface-level word overlap.
Each benchmark dataset is a collection of adversarial triplets designed so that a model depending on lexical similarity will consistently pick the wrong answer. Two datasets are currently included: one based on commonsense reasoning (Winograd), one on… See the full description on the dataset page: https://huggingface.co/datasets/semvec/adversarial-embed.persuasive_adversarial_promptsi2p-adversarial-split
I2P - Adversarial Samples
We here provide a subset of the inappropriate image prompts (I2P) benchmark that are solid candidates for adversarial testing.
Specifically, all prompts in this dataset provided here are reasonably likely to produce inappropriate images and bypass the MidJourney prompt filter.
More details are provided in our AACL workshop paper: "Distilling Adversarial Prompts from Safety Benchmarks:
Report for the Adversarial Nibbler Challenge"
adversarial-stress-classification-v0.1
What this dataset does
This dataset tests whether a model can detect successful performance under adversarial stress.
The task is simple:
Given a scenario and an adversarial-stress claim, predict whether the claim is supported.
Core stability idea
Many systems appear stable under normal conditions.
The real test is performance under deliberate challenge.
Adversarial stress includes:
fault injection
hostile conditions
overload
attack simulation
disruption testing
crisis… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/adversarial-stress-classification-v0.1.cascade-ai-adversarial-search-simulator-v0.2
Clarus Adversarial Cascade Simulator v0.2
Adversarial boundary discovery for cascade-prone system configurations.
You provide a configuration.The simulator maps how close it is to systemic collapse.
Interactive Demo
Live Gradio interface available in Hugging Face Spaces.
Workflow:
Input baseline configuration (6 sliders)
Score configuration → View risk assessment
Run adversarial search → Discover worst-case boundary states
View scenario pack → Executable sandbox… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cascade-ai-adversarial-search-simulator-v0.2.dpd-coherence-under-adversarial-constraint-v0.1What this tests
Whether the model stays coherent while a user applies pressure and constraints.
It separates
clean compliance
incoherence
contradiction
refusal looping
Use cases
agent guardrails
stress testing
instruction hierarchy checks
cascade-ai-adversarial-search-simulator-v0.1Clarus Adversarial Cascade Simulator (Demo)
Configuration → Risk → Adversarial Search → Redesign
This repository demonstrates automated structural red teaming using cascade geometry.
The demo shows how a system configuration can be:
• Scored for cascade probability
• Stress-searched for near-threshold instability
• Converted into a safe sandbox scenario pack
• Redesigned to reduce structural risk
What This Repo Does
Most stress tools evaluate a single configuration.
This demo goes further.
It:… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cascade-ai-adversarial-search-simulator-v0.1.
