datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
multi-turn_jailbreak_attack_datasets
Multi-Turn Jailbreak Attack Datasets
Description
This dataset was created to compare single-turn and multi-turn jailbreak attacks on large language models (LLMs). The primary goal is to take a single harmful prompt and distribute the harm over multiple turns, making each prompt appear harmless in isolation. This approach is compared against traditional single-turn attacks with the complete prompt to understand their relative impacts and failure modes. The key feature of… See the full description on the dataset page: https://huggingface.co/datasets/tom-gibbs/multi-turn_jailbreak_attack_datasets.Visco-Attack
VisCo Attack: Visual Contextual Jailbreak Dataset
📄 arXiv:2507.02844 · 💻 Code – Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection
This dataset contains the adversarial contexts, prompts, and images from the paper: "Visual Contextual Attack: Jailbreaking MLLMs with Image-Driven Context Injection".
⚠️ Content Warning
This dataset contains content that is offensive and/or harmful. It was created for research purposes to study the… See the full description on the dataset page: https://huggingface.co/datasets/miaozq/Visco-Attack.instruction-attack-outputsattackontitan
Bangumi Image Base of Attack On Titan
This is the image base of bangumi Attack On Titan, we detected 76 characters, 14308 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/attackontitan.3MAD-66Kattackdex-paldeaSingle pokemon datasets containing all the attacks (from levelling or TMs) learnable by the relative monster. All the data refer to the Paldea region and they come from the project discussed in https://medium.com/@virtualmartire/i-built-an-algorithm-that-finds-the-optimal-pokemon-team-01ea152824a9.
M-Attack_AdvSamples
M-Attack Adversarial Samples Dataset
This dataset contains 100 adversarial samples generated using M-Attack to perturb the images from the NIPS 2017 Adversarial Attacks and Defenses Competition. This dataset is used in the paper A Frustratingly Simple Yet Highly Effective Attack Baseline: Over 90% Success Rate Against the Strong Black-box Models of GPT-4.5/4o/o1.
Dataset Description
The dataset consists of total 300 adversarial samples organized in three… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI-LLM/M-Attack_AdvSamples.AttackViz
AttackViz
AttackViz is a chart-image dataset for studying correct and misleading data visualizations. It was introduced in the paper ChartAttack: Testing the Vulnerability of LLMs to Malicious Prompting in Chart Generation.
Each example contains a rendered chart image, metadata about the chart and question type, the expected gold answer, a binary label indicating whether the chart is correct or misleading, a misleading-visualization category, and serialized chart annotations.… See the full description on the dataset page: https://huggingface.co/datasets/INSAIT-Institute/AttackViz.ipi_arena_attacks
IPI Arena Attacks
Attack strings from the IPI Arena benchmark for evaluating model robustness to indirect prompt injection (IPI), from How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition.
Dataset
95 attack strings across 28 behaviors, sourced from Qwen (qwen/qwen3-vl-235b-a22b-instruct). These attacks succeeded on open-source models but did not transfer to any closed-source model in the arena.
Each row contains:… See the full description on the dataset page: https://huggingface.co/datasets/sureheremarv/ipi_arena_attacks.3MAD-Tiny-1Katari-demon_attack-dataset
Atari-Demon Attack Dataset
This is a large dataset of 10M video frames and actions collected from the Demon Attack atari environment (Bellemare et al., 2012) in order to train world models.The dataset enables reproducible, large-scale experiments in action-conditioned video prediction. It is meant to be used with Jasmine, our JAX-based world modeling codebase.
Dataset Summary
Environment: Atari Learning Environment
Frames: 10 million
Resolution: 84 × 84… See the full description on the dataset page: https://huggingface.co/datasets/p-doom/atari-demon_attack-dataset.mitre-attack-en
MITRE ATT&CK Enterprise - Complete English Dataset
Comprehensive dataset of the MITRE ATT&CK Enterprise framework on Hugging Face. Data extracted automatically from official STIX 2.1 sources.
Description
This dataset covers the entire MITRE ATT&CK Enterprise framework:
14 tactics with full descriptions
691 techniques and sub-techniques (216 techniques + 475 sub-techniques)
44 mitigations with associated techniques
172 threat groups (APTs) with their known techniques
30… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/mitre-attack-en.Cybersecurity_Attackmitre-attack-synthetic-scenarios
MITRE ATT&CK Synthetic Scenario Logs v3.0
Expanded Dataset: 30 scenarios × 8 events = 240 synthetic events
Axis
Coverage
Environment
endpoint, cloud, SaaS, identity, CI/CD, OT/IoT
Actor Type
external_apt, ransomware, insider, compromised_vendor, careless_admin, automated_threat
Intent
exfiltration, impact, fraud, persistence, reconnaissance, cryptomining, espionage
Detection Source
EDR, IAM, SIEM, DLP, DNS, proxy, cloud_audit, email_gateway, CASB, NDR, PAM, firewall… See the full description on the dataset page: https://huggingface.co/datasets/koushikcs09/mitre-attack-synthetic-scenarios.Replay_attack_mobile
Liveness Detection Replay Dataset (3K+ Attacks, 1.5K People)
iBeta Level 1 Dataset
Liveness detection dataset of Replay attacks performed on Mobile devices. This dataset consists of 1,500 individuals who provided selfies, followed by 3,000 replay display attacks executed across 15 different mobile devices. These attacks are captured from a diverse range of devices, spanning low, medium, and high-end mobile phones, providing extensive variation in screen types, lighting… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/Replay_attack_mobile.ae-signal_processing_attacks_whisper_librispeechattack_data_hfToxicity contail three types of data. 1. from realtoxicty prompt .2 response from gpt3.5 generation as prompt 3. same as 2 but it comes from gpt4
3D_paper_mask_attack_dataset_for_Liveness
Liveness Detection Dataset: 3D Paper Mask Attacks
Paper mask attack dataset for training PAD and liveness detection models against low-cost 3D spoofing — 2,000+ videos on iOS and Android, ISO 30107-3 Level 1/2 attack category
What This Dataset Covers
2,000+ video recordings of 3D paper mask presentation attacks — masks with volumetric elements that simulate facial depth. Captured on iOS and Android devices, ~7 sec per video, with zoom-in/zoom-out phases for active… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/3D_paper_mask_attack_dataset_for_Liveness.audio-hallucination-attack
Audio Hallucination Attacks (AHA)
Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models"
It contains two subsets:
AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs
AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training
Audio Files
The audio files are provided as compressed archives in this repository:
File
Contents
Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.mars-attacks-adam-datasetdemon_attack_pi0_traininstruction-attack-datacl-signal_processing_attacks_whisper_librispeech
Dataset Card for "cl-signal_processing_attacks_large"
More Information needed
prompt-injection-attack-datasetprint-attack-dataset
Liveness Detection Dataset: Photo Print attack dataset (3K individuals+)
What Is a Print Attack?
A print attack is a 2D presentation attack vector against face recognition and liveness detection systems, where an attacker presents a printed photo of a real person's face to a camera to deceive biometric authentication. Print attacks are the most common and accessible spoofing technique in face anti-spoofing research and represent the entry-level attack class tested in… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/print-attack-dataset.video-framesxai-attack-detection-cifar10
XAI Attack Detection — CIFAR-10 PGD
This private research dataset contains balanced, paired clean and adversarial images for
studying whether an attack can be detected from a classifier explanation map.
Dataset construction
The source is the CIFAR-10 test split. A fine-tuned OpenCLIP ViT-B/16 classifies each
clean image. Clean-correct examples are attacked with untargeted L-infinity PGD using
epsilon 8/255, step size 2/255, 10 steps, and deterministic random… See the full description on the dataset page: https://huggingface.co/datasets/nimaeb/xai-attack-detection-cifar10.Chat-Models-Backdoor-AttackingHere are the data for the paper "Exploring Backdoor Attacks on Chat Models"[paper],
including both the chat data and instructional data. The structure of the whole data is shown
below:
Chat_Data
|-- Poisoned_dataset
| |-- BenignScn_MaliciousScn
| | |-- General_Harmless_Data_10K.json
| | |-- Helpful_Data_10K.json
| | |-- Multi-TS_Poisoned_Data_2K.json
| | |-- Single-TS_Harmless_Data_2K.json
| | |-- Poisoned_Data_24K.json
| |--… See the full description on the dataset page: https://huggingface.co/datasets/luckychao/Chat-Models-Backdoor-Attacking.vulnerability-attack-techniques
vulnerability-attack-techniques
This dataset maps 1,207 CVEs to MITRE ATT&CK (Enterprise) techniques, joining
hand-curated mappings from the MITRE Center for Threat-Informed Defense (CTID)
with vulnerability descriptions from
CIRCL/vulnerability-scores.
It is intended for training and evaluating models that suggest candidate ATT&CK
techniques from a vulnerability description: CVSS tells you how bad a
vulnerability is, CWE what kind of flaw it is — ATT&CK tells defenders what… See the full description on the dataset page: https://huggingface.co/datasets/CIRCL/vulnerability-attack-techniques.Display_replay_attacks
Face Anti Spoofing Replay Dataset
iBeta Level 1 Dataset
Liveness Detection: Replay attacks. 5,000+ videos of display replay monitor attacks 12+ sec and real photos. The attacks provide diversity of lighting, devices, and screens
Full version of dataset is availible for commercial usage - leave a request on our website Axon Labs to purchase the dataset 💰
Left: Real selfie; Right: Display attack
Left: Real selfie; Right: Display attack
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/Display_replay_attacks.
