datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
successful_adversarial_prompts
Citation
If you use this dataset, please cite the associated paper:
@article{chugh2026recap,
title = {RECAP: A Resource-Efficient Method for Adversarial Prompting in Large Language Models},
author = {Chugh, Rishit},
journal = {arXiv preprint arXiv:2601.15331},
year = {2026},
url = {https://arxiv.org/abs/2601.15331}
}
sffop_1706381144_410msft_relabel_pythia6.9b_logprobs_prefix_chosenNemotron-RL-Instruction-Following-Adversarial-v1-prompt-only
Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Instruction-Following-Adversarial-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Instruction-Following-Adversarial-v1-prompt-only.POD-DeepONet-Adversarial-Activationslang-adversarial-inference-01
Inference-Layer Security: Defending Against Adversarial Inference and Infrastructure Abuse
An exploration on how to secure Model-as-a-Service infrastructure at the inference layer against runtime exploits, ranging from automated bot farms to adversarial extraction and agentic misuse
More details about the project: https://www.daoist.dev/posts/adversarial-inference-security-1
nq_retrieved_adversarial_passageadversarial-vision-transformerssummarize_from_feedback_tldr_3_filtered_oai_preprocessing_1706381144_propprefixsummarize_from_feedback_oai_preprocessing_1706381144_410msft_relabel_pythia6.9b_logprobsduel-adversarial-logs
DUEL Adversarial Security Dataset
Dataset Description
This dataset contains synthetic adversarial telemetry generated by the DUEL framework
(Dual Unified Evasion Loop) — an adversarial LLM security research framework where an
Attacker agent and a Defender agent battle across MITRE ATT&CK and OWASP LLM Top 10
techniques against real Microsoft Sentinel schemas.
Every record is a single synthetic log entry from one round of the adversarial loop,
labelled as evaded (the… See the full description on the dataset page: https://huggingface.co/datasets/0xDanielSec/duel-adversarial-logs.repro-consistent-adversarial-attacks-traces
Agent traces
Agent sessions published from a Trackio Logbook.
sffop_1706381144_410msft_relabel_pythia6.9b_logprobs_cond3emojieallprefixafrica-ai-health-adversarial
African AI Health Adversarial Dataset | Africa (original)
Size category: 10K<n<100K - Formats: parquet, optimized-parquet - Sector: health - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Health datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ai-health-adversarial.AdversarialArena_Nova_AI_Challenge_Trusted_AI_Dataset
Adversarial Arena: Trusted AI Challenge Dataset
Dataset Description
This dataset contains multi-turn adversarial conversations generated through the Adversarial Arena framework, an interactive competition where attacker bots attempt to elicit unsafe code or cyberattack assistance from defender bots. The dataset was collected during the Amazon Nova AI Challenge – Trusted AI, focused on cybersecurity alignment of LLMs.
Papers:
Adversarial Arena: Crowdsourcing Data… See the full description on the dataset page: https://huggingface.co/datasets/amazon-agi/AdversarialArena_Nova_AI_Challenge_Trusted_AI_Dataset.i2p-adversarial-split
I2P - Adversarial Samples
We here provide a subset of the inappropriate image prompts (I2P) benchmark that are solid candidates for adversarial testing.
Specifically, all prompts in this dataset provided here are reasonably likely to produce inappropriate images and bypass the MidJourney prompt filter.
More details are provided in our AACL workshop paper: "Distilling Adversarial Prompts from Safety Benchmarks:
Report for the Adversarial Nibbler Challenge"
sffop_1706381144_410msft_relabel_pythia6.9b_logprobs_cond3emojiepropallprefixsffop_1706381144_410msft_relabel_pythia6.9b_logprobs_cond3emojiebothlm-eval-EleutherAI_tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748
Dataset Card for Evaluation run of EleutherAI/tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748
Dataset automatically created during the evaluation run of model EleutherAI/tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748
The dataset is composed of 2 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 11 run(s). Each run can be found as a specific split in each configuration, the split being… See the full description on the dataset page: https://huggingface.co/datasets/EleutherAI/lm-eval-EleutherAI_tampered-deep-ignorance-random-init-fp-adversarial-20251104_051748.nq_retrieved_adversarial_sentence_simsffop_1706381144_410msft_relabel_pythia6.9b_logprobs_cond3emojiepropprefixsffop_1706381144_410msft_relabel_pythia6.9b_logprobs_cond3emojieprefixppo_pythia410m_tldr6.9b_rm410mdata_mergedsft_propprefix_eval-datasetevent-adversarial-gptppo_pythia410m_tldr6.9b_rm410mdata_mergedsft_prefix_nokl_full_eval-datasetsffop_1706381144_410msft_relabel_pythia6.9b_3emojieprefix_randomizerloo_pythia410m_tldr6.9b_rm410mdata_mergedsft_prefix_eval-datasetadversarial-vision-transformers-robustnesssffop_1706381144_410msft_relabel_pythia6.9b_logprobs_cond3emojiesuffix410M-sft-tldr-eval-datasetadversarial-mnist
MNIST with Adversarial Examples
This dataset contains MNIST images with both normal and adversarial examples.
The dataset includes:
Original MNIST digit images (28x28 pixels, flattened to 784 features)
Adversarial examples generated from the original images
Labels for digit classification (0-9)
Binary flag indicating whether each sample is adversarial
Features:
label: Digit class (0-9)
pixels 0-783: Flattened 28x28 grayscale pixel values
is_adversarial: Binary flag (0 = normal, 1… See the full description on the dataset page: https://huggingface.co/datasets/wambosec/adversarial-mnist.
