datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ipi_arena_attacks
IPI Arena Attacks
Attack strings from the IPI Arena benchmark for evaluating model robustness to indirect prompt injection (IPI), from How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition.
Dataset
95 attack strings across 28 behaviors, sourced from Qwen (qwen/qwen3-vl-235b-a22b-instruct). These attacks succeeded on open-source models but did not transfer to any closed-source model in the arena.
Each row contains:… See the full description on the dataset page: https://huggingface.co/datasets/sureheremarv/ipi_arena_attacks.NADW-network-attacks-dataset
Network Traffic Dataset for Anomaly Detection
Overview
This project presents a comprehensive network traffic dataset used for training AI models for anomaly detection in cybersecurity. The dataset was collected using Wireshark and includes both normal network traffic and various types of simulated network attacks. These attacks cover a wide range of common cybersecurity threats, providing an ideal resource for training systems to detect and respond to real-time network… See the full description on the dataset page: https://huggingface.co/datasets/onurkya7/NADW-network-attacks-dataset.attacks-monthly
attacks-monthly
Dataset Description
This dataset contains cybersecurity events collected from honeypot infrastructure.
The data has been processed and feature-engineered for machine learning applications in threat detection and security analytics.
Feature Categories
Network Features
Connection flow statistics (bytes, packets, duration)
Protocol-specific metrics
Geographic information
IP reputation data
Behavioral Features
Session… See the full description on the dataset page: https://huggingface.co/datasets/pyToshka/attacks-monthly.security-attacks-MITREweb-attacks-longllm-prompt-injection-attacks
Prompt Injection Mechanisms Dataset
Overview
A 55,000-sample multi-label dataset for prompt injection detection in large language models.
Labels
- BENIGN
- JAILBREAK
- INSTRUCTION_OVERRIDE
- ROLE_HIJACK
- DATA_EXFILTRATION
Format
The dataset is provided in Apache Parquet format with train/validation splits.
Construction
The dataset was created by merging multiple public prompt-injection datasets and
re-annotating them using a… See the full description on the dataset page: https://huggingface.co/datasets/Smooth-3/llm-prompt-injection-attacks.web-attacksMitre_Attacks_Framework_Dataset
MITRE ATT&CK Enterprise Dataset
Overview
This dataset provides a comprehensive collection of MITRE ATT&CK Enterprise techniques (v14.1) in JSONL format, designed for cybersecurity professionals, red teams, and threat hunters.
Each entry maps to a specific ATT&CK technique, including its ID, name, description, real-world example, and source.
The dataset is structured for seamless integration into security tools such as SIEMs, threat intelligence platforms, or custom red… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Mitre_Attacks_Framework_Dataset.llm-prompt-injection-attacks
Prompt Injection Mechanisms Dataset
Overview
A 55,000-sample multi-label dataset for prompt injection detection in large language models.
Labels
- BENIGN
- JAILBREAK
- INSTRUCTION_OVERRIDE
- ROLE_HIJACK
- DATA_EXFILTRATION
Format
The dataset is provided in Apache Parquet format with train/validation splits.
Construction
The dataset was created by merging multiple public prompt-injection datasets and
re-annotating them using a… See the full description on the dataset page: https://huggingface.co/datasets/cyberec/llm-prompt-injection-attacks.attacks-daily
attacks-daily
Dataset Description
This dataset contains cybersecurity events collected from honeypot infrastructure.
The data has been processed and feature-engineered for machine learning applications in threat detection and security analytics.
Feature Categories
Network Features
Connection flow statistics (bytes, packets, duration)
Protocol-specific metrics
Geographic information
IP reputation data
Behavioral Features
Session… See the full description on the dataset page: https://huggingface.co/datasets/pyToshka/attacks-daily.geometry-of-harmfulness-in-multi-turn-attacks
Geometry of Harmfulness — Multi-Turn Attack Conversations
Raw multi-turn attack conversations accompanying the paper
The Geometry of Harmfulness in Multi-Turn Attacks. These are the conversations
from which the paper's hidden-state representations are extracted; the analysis
code lives in the companion repository.
Conversations were generated by running three multi-turn attack frameworks —
Crescendo, ActorAttack, and X-Teaming (attacker & judge: GPT-4o) —
against three… See the full description on the dataset page: https://huggingface.co/datasets/yelyzavetahusieva/geometry-of-harmfulness-in-multi-turn-attacks.repro-consistent-adversarial-attacks-traces
Agent traces
Agent sessions published from a Trackio Logbook.
agent-social-eng-attacks-sample
Dataset Description
Small sample of attack vectors involving a user attempt to exfiltrate information from a user, by means of social engineering technqiues.
Dataset Details
Created by: Always Further
License: CC BY 4.0
Language(s): [English
Dataset Size: 737
Data Splits
[train]
Usage
from datasets import load_dataset
dataset = load_dataset("always-further/agent-social-eng-attacks-sample/")
Limitations and Bias
Use… See the full description on the dataset page: https://huggingface.co/datasets/nolabs/agent-social-eng-attacks-sample.voice-jailbreak-attacks
Voice Jailbreak Attacks — synthesised audio
Pre-synthesised speech for Voice Jailbreak Attacks Against GPT-4o
(Shen et al., 2024), spoken with OpenAI
text-to-speech (tts-1, voice fable) from the forbidden questions and text
jailbreak templates published in the paper's
repository.
Content warning: this dataset contains spoken instructions for fraud,
hate speech, illegal activity, physical harm, pornography, and privacy
violations, and jailbreak templates written to elicit… See the full description on the dataset page: https://huggingface.co/datasets/ymerkli/voice-jailbreak-attacks.africa-terrorism-attacks-in-kenya
Terrorism Attacks in Kenya | Africa (original)
Size category: n<1K - Formats: parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This Dataset Covers
Public datasets help analysts inspect structured… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-terrorism-attacks-in-kenya.llm-adaptive-attacksafrica-south-sudan-south-sudan-ssd-attacks-on-aid-operations-education-health-ebed880f
South Sudan Ssd Attacks On Aid Operations Education Health | Africa (Insecurity Insight)
16,432 rows - 1 Africa country/area - 2024-2026 - 26 indicators - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 16,432 rows from Insecurity Insight, covering South Sudan Ssd Attacks On Aid Operations Education Health. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-south-sudan-south-sudan-ssd-attacks-on-aid-operations-education-health-ebed880f.africa-somalia-somalia-som-attacks-on-aid-operations-education-food-and-w-cae90e93
Somalia Som Attacks On Aid Operations Education Food and W | Africa (Insecurity Insight)
36 rows - 1 Africa country/area - 2023-2025 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 36 rows from Insecurity Insight, covering Somalia Som Attacks On Aid Operations Education Food and W. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-somalia-somalia-som-attacks-on-aid-operations-education-food-and-w-cae90e93.africa-somalia-attacks-on-public-health-related-programmes-data-2200d2e0
Attacks On Public Health Related Programmes Data | Africa (Insecurity Insight)
1,180 rows - 1 Africa country/area - 2020 - source table - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 1,180 rows from Insecurity Insight, covering Attacks On Public Health Related Programmes Data. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-somalia-attacks-on-public-health-related-programmes-data-2200d2e0.africa-drc-attacks-on-health-care-data-00869950
Attacks on Health Care Data | Africa (DRC official open data)
27,168 rows - 1 Africa country - 2026 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official XLSX resource from DRC as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: Attacks on Health Care Data
Publisher: Insecurity Insight
Resource:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-drc-attacks-on-health-care-data-00869950.supply-chain-attacks-en
Software Supply Chain Attacks - English Dataset
The reference English dataset on software supply chain attacks.
This dataset contains detailed and structured information on supply chain attack vectors, real-world incidents, defense mechanisms, and 80 English Q&A pairs. It is designed for training, awareness, and fine-tuning language models specialized in cybersecurity.
Dataset Contents
Type
Count
Description
Attack Vectors
~30
Vectors organized by category… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/supply-chain-attacks-en.africa-mali-attacks-on-public-health-related-programmes-data-2200d2e0
Attacks on Public Health-Related Programmes Data | Africa (Mali official open data)
1,180 rows - 1 Africa country - 2020 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official XLSX resource from Mali as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: Attacks on Public Health-Related Programmes… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-mali-attacks-on-public-health-related-programmes-data-2200d2e0.asia-owid-deaths-from-terrorist-attacks-by-method
Deaths From Terrorist Attacks By Method | Asia (Our World in Data)
🌏 2,178 observations · 46 Asia countries · 1970–2021 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 2,178 observations of Deaths From Terrorist Attacks By Method data across 46 Asia countries, spanning 1970–2021.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Deaths From Terrorist Attacks By Method… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-deaths-from-terrorist-attacks-by-method.africa-somalia-attacks-on-public-health-related-programmes-data-5735f657
Attacks On Public Health Related Programmes Data | Africa (Insecurity Insight)
44 rows - 1 Africa country/area - 2026 - 8 indicators - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 44 rows from Insecurity Insight, covering Attacks On Public Health Related Programmes Data. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-somalia-attacks-on-public-health-related-programmes-data-5735f657.africa-drc-democratic-republic-of-the-congo-cod-attacks-on-aid-operat-4de6ac3b
Democratic Republic of the Congo (COD): Attacks on Aid Operations, Education, Food and Water Systems, Health Care and IDP/Refugee Camps, and Conflict-Related or Political-Related Sexual Violence and Explosive Weapons Incident Data | Africa (DRC official open data)
9,334 rows - 1 Africa country - 2024-2026 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official XLSX resource from DRC as
ML-ready Parquet. The source file is the provenance… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-drc-democratic-republic-of-the-congo-cod-attacks-on-aid-operat-4de6ac3b.attacks-weekly
attacks-weekly
Dataset Description
This dataset contains cybersecurity events collected from honeypot infrastructure.
The data has been processed and feature-engineered for machine learning applications in threat detection and security analytics.
Feature Categories
Network Features
Connection flow statistics (bytes, packets, duration)
Protocol-specific metrics
Geographic information
IP reputation data
Behavioral Features
Session… See the full description on the dataset page: https://huggingface.co/datasets/pyToshka/attacks-weekly.africa-ai-llm-attacks
AI-Powered & LLM-Assisted Attacks (Africa) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet, optimized-parquet - Sector: governance_security - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.
What This… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-ai-llm-attacks.africa-somalia-attacks-on-health-care-data-5f769d12
Attacks On Health Care Data | Africa (Insecurity Insight)
303,728 rows - 1 Africa country/area - 2024-2026 - 16 indicators - Engineered by Electric Sheep Africa
TL;DR
This dataset contains 303,728 rows from Insecurity Insight, covering Attacks On Health Care Data. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.
What This Dataset Measures
Health datasets help… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-somalia-attacks-on-health-care-data-5f769d12.africa-mali-mali-mli-attacks-on-aid-operations-education-food-and-wate-71ab805e
Mali (MLI): Attacks on Aid Operations, Education, Food and Water Systems and Health Care | Africa (Mali official open data)
2,006 rows - 1 Africa country - 2024-2026 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official XLSX resource from Mali as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source:… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-mali-mali-mli-attacks-on-aid-operations-education-food-and-wate-71ab805e.africa-drc-attacks-on-public-health-related-programmes-data-c9b6cca6
Attacks on Public Health-Related Programmes Data | Africa (DRC official open data)
493 rows - 1 Africa country - 2018-2020 - Repackaged by Electric Sheep Africa
TL;DR
This dataset packages one official XLSX resource from DRC as
ML-ready Parquet. The source file is the provenance boundary; all usable
indicators or tabular columns from the resource stay together in this repo.
About the source
Source: Attacks on Public Health-Related Programmes… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-drc-attacks-on-public-health-related-programmes-data-c9b6cca6.
