datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ipi_arena_attacks
IPI Arena Attacks
Attack strings from the IPI Arena benchmark for evaluating model robustness to indirect prompt injection (IPI), from How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition.
Dataset
95 attack strings across 28 behaviors, sourced from Qwen (qwen/qwen3-vl-235b-a22b-instruct). These attacks succeeded on open-source models but did not transfer to any closed-source model in the arena.
Each row contains:… See the full description on the dataset page: https://huggingface.co/datasets/sureheremarv/ipi_arena_attacks.ae-signal_processing_attacks_whisper_librispeechcl-signal_processing_attacks_whisper_librispeech
Dataset Card for "cl-signal_processing_attacks_large"
More Information needed
mars-attacks-adam-datasetDisplay_replay_attacks
Face Anti Spoofing Replay Dataset
iBeta Level 1 Dataset
Liveness Detection: Replay attacks. 5,000+ videos of display replay monitor attacks 12+ sec and real photos. The attacks provide diversity of lighting, devices, and screens
Full version of dataset is availible for commercial usage - leave a request on our website Axon Labs to purchase the dataset 💰
Left: Real selfie; Right: Display attack
Left: Real selfie; Right: Display attack
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/Display_replay_attacks.NADW-network-attacks-dataset
Network Traffic Dataset for Anomaly Detection
Overview
This project presents a comprehensive network traffic dataset used for training AI models for anomaly detection in cybersecurity. The dataset was collected using Wireshark and includes both normal network traffic and various types of simulated network attacks. These attacks cover a wide range of common cybersecurity threats, providing an ideal resource for training systems to detect and respond to real-time network… See the full description on the dataset page: https://huggingface.co/datasets/onurkya7/NADW-network-attacks-dataset.low_quality_webcam_video_attacksThe dataset includes live-recorded Anti-Spoofing videos from around the world,
captured via low-quality webcams with resolutions like QVGA, QQVGA and QCIF.2d-masks-pad-attacks
2D Masks with Eyeholes Attacks
The dataset comprises 11,200+ videos of people wearing of holding 2D printed masks with eyeholes captured using 5 different devices. This extensive collection is designed for research in presentation attacks, focusing on various detection methods, primarily aimed at meeting the requirements for iBeta Level 1 & 2 certification. Specifically engineered to challenge facial recognition and enhance spoofing detection techniques.
By utilizing this… See the full description on the dataset page: https://huggingface.co/datasets/UniDataPro/2d-masks-pad-attacks.attacks-monthly
attacks-monthly
Dataset Description
This dataset contains cybersecurity events collected from honeypot infrastructure.
The data has been processed and feature-engineered for machine learning applications in threat detection and security analytics.
Feature Categories
Network Features
Connection flow statistics (bytes, packets, duration)
Protocol-specific metrics
Geographic information
IP reputation data
Behavioral Features
Session… See the full description on the dataset page: https://huggingface.co/datasets/pyToshka/attacks-monthly.MacBook-Attacks-DatasetThe dataset consists of videos of replay attacks played on different
models of MacBooks. The dataset solves tasks in the field of anti-spoofing and
it is useful for buisness and safety systems.
The dataset includes: **replay attacks** - videos of real people played on
a computer and filmed on the phone.web-attacks-longsecurity-attacks-MITREmcl-signal_processing_attacks_whisper_librispeech
Dataset Card for "mcl-signal_processing_attacks_large"
More Information needed
ad-attacks-en
Active Directory Attacks - Complete English Dataset
Comprehensive dataset of Active Directory attacks on Hugging Face. Complete reference for offensive and defensive security in AD environments.
Description
This dataset covers all known Active Directory attack techniques:
46 attacks documented with detailed descriptions, prerequisites, tools, detection and mitigation
33 AD pentest tools (Mimikatz, Impacket, BloodHound, Rubeus, etc.)
30 detection rules in Sigma format… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ad-attacks-en.high_quality_webcam_video_attacksThe dataset includes live-recorded Anti-Spoofing videos from around the world,
captured via **high-quality** webcams with Full HD resolution and above.llm-prompt-injection-attacks
Prompt Injection Mechanisms Dataset
Overview
A 55,000-sample multi-label dataset for prompt injection detection in large language models.
Labels
- BENIGN
- JAILBREAK
- INSTRUCTION_OVERRIDE
- ROLE_HIJACK
- DATA_EXFILTRATION
Format
The dataset is provided in Apache Parquet format with train/validation splits.
Construction
The dataset was created by merging multiple public prompt-injection datasets and
re-annotating them using a… See the full description on the dataset page: https://huggingface.co/datasets/Smooth-3/llm-prompt-injection-attacks.web-attacksMitre_Attacks_Framework_Dataset
MITRE ATT&CK Enterprise Dataset
Overview
This dataset provides a comprehensive collection of MITRE ATT&CK Enterprise techniques (v14.1) in JSONL format, designed for cybersecurity professionals, red teams, and threat hunters.
Each entry maps to a specific ATT&CK technique, including its ID, name, description, real-world example, and source.
The dataset is structured for seamless integration into security tools such as SIEMs, threat intelligence platforms, or custom red… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Mitre_Attacks_Framework_Dataset.monitors-replay-attacks-datasetThe dataset consists of videos of replay attacks played on different models of
computers. The dataset solves tasks in the field of anti-spoofing and it is
useful for buisness and safety systems.
The dataset includes: **replay attacks** - videos of real people played
on a computer and filmed on the phone.silicone-masks-biometric-attacksThe dataset consists of videos of individuals and attacks with printed 2D masks and
silicone masks . Videos are filmed in different lightning conditions (*in a dark room,
daylight, light room and nightlight*). Dataset includes videos of people with different
attributes (*glasses, mask, hat, hood, wigs and mustaches for men*).printed-2d-masks-with-holes-for-eyes-attacksThe dataset consists of selfies of people and videos of them wearing a printed
2d mask with their face. The dataset solves tasks in the field of anti-spoofing
and it is useful for buisness and safety systems.
The dataset includes: **attacks** - videos of people wearing printed portraits
of themselves with cut-out eyes.llm-prompt-injection-attacks
Prompt Injection Mechanisms Dataset
Overview
A 55,000-sample multi-label dataset for prompt injection detection in large language models.
Labels
- BENIGN
- JAILBREAK
- INSTRUCTION_OVERRIDE
- ROLE_HIJACK
- DATA_EXFILTRATION
Format
The dataset is provided in Apache Parquet format with train/validation splits.
Construction
The dataset was created by merging multiple public prompt-injection datasets and
re-annotating them using a… See the full description on the dataset page: https://huggingface.co/datasets/cyberec/llm-prompt-injection-attacks.printed_photos_attacksThe dataset consists of 40,000 videos and selfies with unique people. 15,000
attack replays from 4,000 unique devices. 10,000 attacks with A4 printouts and
10,000 attacks with cut-out printouts.Muli-Generationtion-attacksWrapped_3D_Attacks
Wrapped 3D Attacks Dataset
Full version of dataset is availible for commercial usage - leave a request on our website Axon Labs to purchase the dataset 💰
Introduction
This dataset is designed to enhance Liveness Detection models by simulating Wrapped 3D Attacks — a more advanced version of 3D Print Attacks, where facial prints include 3D elements and additional attributes. It is particularly useful for iBeta Level 2 certification and anti-spoofing model… See the full description on the dataset page: https://huggingface.co/datasets/AxonData/Wrapped_3D_Attacks.attacks-daily
attacks-daily
Dataset Description
This dataset contains cybersecurity events collected from honeypot infrastructure.
The data has been processed and feature-engineered for machine learning applications in threat detection and security analytics.
Feature Categories
Network Features
Connection flow statistics (bytes, packets, duration)
Protocol-specific metrics
Geographic information
IP reputation data
Behavioral Features
Session… See the full description on the dataset page: https://huggingface.co/datasets/pyToshka/attacks-daily.geometry-of-harmfulness-in-multi-turn-attacks
Geometry of Harmfulness — Multi-Turn Attack Conversations
Raw multi-turn attack conversations accompanying the paper
The Geometry of Harmfulness in Multi-Turn Attacks. These are the conversations
from which the paper's hidden-state representations are extracted; the analysis
code lives in the companion repository.
Conversations were generated by running three multi-turn attack frameworks —
Crescendo, ActorAttack, and X-Teaming (attacker & judge: GPT-4o) —
against three… See the full description on the dataset page: https://huggingface.co/datasets/yelyzavetahusieva/geometry-of-harmfulness-in-multi-turn-attacks.attacks-with-2d-printed-masks-of-indian-people
Attacks with 2D Printed Masks of Indian People - Biometric Attack Dataset
The dataset consists of videos of individuals wearing printed 2D masks of different kinds and directly looking at the camera. Videos are filmed in different lightning conditions and in different places (indoors, outdoors). Each video in the dataset has an approximate duration of 3-4 seconds.
The similar dataset that includes all ethnicities - Printed 2D Masks Attacks Dataset
Types of… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/attacks-with-2d-printed-masks-of-indian-people.biometric-attacks-in-different-lighting-conditions
Biometric Attack Dataset - Different Lighting Conditions Dataset
The liveness detection dataset consists of videos of individuals and attacks with photos shown in the monitor . Videos are filmed in different lightning conditions (in a dark room, daylight, light room and nightlight) and in different places (indoors, outdoors). Each video in the dataset has an approximate duration of 20 seconds.
The dataset is created on the basis of iBeta Level 1 Dataset… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/biometric-attacks-in-different-lighting-conditions.NADW-network-attacks-dataset
Network Traffic Dataset for Anomaly Detection
Overview
This project presents a comprehensive network traffic dataset used for training AI models for anomaly detection in cybersecurity. The dataset was collected using Wireshark and includes both normal network traffic and various types of simulated network attacks. These attacks cover a wide range of common cybersecurity threats, providing an ideal resource for training systems to detect and respond to real-time… See the full description on the dataset page: https://huggingface.co/datasets/aadankan/NADW-network-attacks-dataset.
