datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ipi_arena_attacks
IPI Arena Attacks
Attack strings from the IPI Arena benchmark for evaluating model robustness to indirect prompt injection (IPI), from How Vulnerable Are AI Agents to Indirect Prompt Injections? Insights from a Large-Scale Public Competition.
Dataset
95 attack strings across 28 behaviors, sourced from Qwen (qwen/qwen3-vl-235b-a22b-instruct). These attacks succeeded on open-source models but did not transfer to any closed-source model in the arena.
Each row contains:… See the full description on the dataset page: https://huggingface.co/datasets/sureheremarv/ipi_arena_attacks.mitre-attack-synthetic-scenarios
MITRE ATT&CK Synthetic Scenario Logs v3.0
Expanded Dataset: 30 scenarios × 8 events = 240 synthetic events
Axis
Coverage
Environment
endpoint, cloud, SaaS, identity, CI/CD, OT/IoT
Actor Type
external_apt, ransomware, insider, compromised_vendor, careless_admin, automated_threat
Intent
exfiltration, impact, fraud, persistence, reconnaissance, cryptomining, espionage
Detection Source
EDR, IAM, SIEM, DLP, DNS, proxy, cloud_audit, email_gateway, CASB, NDR, PAM, firewall… See the full description on the dataset page: https://huggingface.co/datasets/koushikcs09/mitre-attack-synthetic-scenarios.attack_data_hfToxicity contail three types of data. 1. from realtoxicty prompt .2 response from gpt3.5 generation as prompt 3. same as 2 but it comes from gpt4
audio-hallucination-attack
Audio Hallucination Attacks (AHA)
Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models"
It contains two subsets:
AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs
AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training
Audio Files
The audio files are provided as compressed archives in this repository:
File
Contents
Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.box-attackshell-attack-evolution-dataset
Shell Honeypot Attack Request–Response Dataset
A standardized, MITRE ATT&CK–annotated dataset of post-login shell
attacks captured by Cowrie SSH/Telnet
honeypots across two collection periods — 2021–2022 and 2024. It pairs
attacker shell commands with real captured system responses, enabling both
longitudinal threat analysis and the training/evaluation of AI-driven honeypots.
This is the open-source release accompanying the paper “Unveiling Evolving
Threats: A Data Analysis… See the full description on the dataset page: https://huggingface.co/datasets/zyw-286/shell-attack-evolution-dataset.security-attacks-MITREgeo-injection-rag-attack-data
Can It Reach the Generator? Investigating the Survival of GEO Prompt-Injection Attacks in Realistic RAG Settings
This dataset contains the prompt-injection attack data presented in the paper Can It Reach the Generator? Investigating the Survival of Prompt-Injection Attacks in Realistic RAG Settings.
The dataset is used to… See the full description on the dataset page: https://huggingface.co/datasets/Euanyu/geo-injection-rag-attack-data.2026-08-28-t2-9284-attack716-train
Attack-variant robustness mixture (9,284 + 716)
field
value
experiment
Training mixture for the attack-variant robustness arm. 179 difficult-advice scenarios x four user-prompt framings that all press for the same norm-violating shortcut (original, hidden, incremental, authority); the assistant reply is held constant across a scenario's four framings (the correct refusal). Plus the same 9,284 Table2 rows. Trains the model to hold its line when the ask is reframed.… See the full description on the dataset page: https://huggingface.co/datasets/matboz/2026-08-28-t2-9284-attack716-train.Mitre_Attacks_Framework_Dataset
MITRE ATT&CK Enterprise Dataset
Overview
This dataset provides a comprehensive collection of MITRE ATT&CK Enterprise techniques (v14.1) in JSONL format, designed for cybersecurity professionals, red teams, and threat hunters.
Each entry maps to a specific ATT&CK technique, including its ID, name, description, real-world example, and source.
The dataset is structured for seamless integration into security tools such as SIEMs, threat intelligence platforms, or custom red… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Mitre_Attacks_Framework_Dataset.CoER-Attacker-SFT
CoER Attacker SFT
Project page · Paper · Code
Stage 1 supervision for initializing the adaptive attacker. Successful conversations retain all attacker turns, including earlier attempts that provide context for later adaptation.
Contents
Split
File
Size
train
train.jsonl
3,995 conversations
The corpus contains 11,655 assistant/attacker turns. Preserve all assistant-turn supervision; do not reduce a conversation to its final payload.
Load… See the full description on the dataset page: https://huggingface.co/datasets/Z-Edgar/CoER-Attacker-SFT.tram-attack-multilabel-clean
TRAM ATT&CK Multi-Label (cleaned, with leak-free splits)
Sentence-level multi-label mapping of cyber threat intelligence prose to MITRE
ATT&CK technique IDs. Derived from MITRE CTID's
TRAM corpus,
deduplicated and republished with two split schemes so that leakage can be
measured rather than assumed.
Everything here is regenerated by python scripts/01_build_dataset.py. No row
was edited by hand.
Why this exists
The upstream corpus contains 19,178 sentences drawn… See the full description on the dataset page: https://huggingface.co/datasets/ctokx/tram-attack-multilabel-clean.SafeDecoding-Attackers
Dataset Details
This dataset contains attack prompts generated from GCG, AutoDAN, PAIR, and DeepInception for research use ONLY.
Dataset Sources
Repository: https://github.com/uw-nsl/SafeDecoding
Paper: https://arxiv.org/abs/2402.08983
Mitre-ATTACK-reasoning-datasetllm-routing-attack-data
MauroPello/llm-routing-attack-data
This dataset contains the JSONL splits used for LLM routing attack experiments.
The Hub exposes the full dataset as the default configuration and the smaller sample as the reduced configuration when both are present.
Files
File
Rows
Size
full/train.jsonl
68687
63.1 MB
full/val.jsonl
14721
13.5 MB
full/test.jsonl
14721
13.5 MB
reduced/train.jsonl
4382
4.1 MB
reduced/val.jsonl
941
906.4 KB
reduced/test.jsonl… See the full description on the dataset page: https://huggingface.co/datasets/MauroPello/llm-routing-attack-data.repro-consistent-adversarial-attacks-traces
Agent traces
Agent sessions published from a Trackio Logbook.
mitre-attack-groups
RelayShield MITRE ATT&CK Group-Technique Mapping
A structured slice of MITRE ATT&CK Enterprise data: 189 named threat actor groups, each mapped to its associated ATT&CK techniques and software, with descriptions and source citations.
This is a cleaned, machine-readable export of MITRE's public STIX bundle — useful if you want group→technique mappings without parsing STIX yourself.
Fields
Field
Type
Description
group_id
string
MITRE ATT&CK group ID (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/relayshieldadmin/mitre-attack-groups.attack-mitigation-corpus-v1
Attack-Mitigation-Corpus v1
One JSONL file of 123 security records. Each record is either a described attack
technique or a mitigation, and each attack record is paired with a mitigation record
covering a related defensive concern.
Measured composition
metric
value
command
rows
123
wc -l < attack_mitigation_corpus.jsonl
kind == "attack"
106
python3 -c "import json,collections;print(collections.Counter(json.loads(l)['kind'] for l in… See the full description on the dataset page: https://huggingface.co/datasets/CatQualia/attack-mitigation-corpus-v1.agentshield-attack-scenarios
AgentShield Attack Scenarios
100 structured attack scenarios for evaluating security of agentic AI systems operating in high-risk domains (biological research). Developed alongside the AgentShield detection pipeline.
Dataset Description
Each scenario is a multi-turn attack sequence targeting one of four attack surfaces identified through STRIDE threat modeling of a BioTeam-AI multi-agent system:
ID
Attack Surface
AS-001
Agent-to-Agent Communication… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/agentshield-attack-scenarios.Mitre_Attacks_Framework_Dataset
MITRE ATT&CK Enterprise Dataset
Overview
This dataset provides a comprehensive collection of MITRE ATT&CK Enterprise techniques (v14.1) in JSONL format, designed for cybersecurity professionals, red teams, and threat hunters.
Each entry maps to a specific ATT&CK technique, including its ID, name, description, real-world example, and source.
The dataset is structured for seamless integration into security tools such as SIEMs, threat intelligence platforms, or custom red… See the full description on the dataset page: https://huggingface.co/datasets/JR87/Mitre_Attacks_Framework_Dataset.cti-attack-synthetic-augmentation
CTI ATT&CK synthetic augmentation
Synthetic training sentences labeled with MITRE ATT&CK technique IDs, built to
augment the training set of a defensive, sentence-level ATT&CK classifier.
Multi-label, 49 techniques.
These sentences are machine-generated. They are not real threat reports. They
exist to add training signal, especially for the rare techniques the real corpus
barely covers. They are for training only, and were never used to evaluate any
model.
What is in… See the full description on the dataset page: https://huggingface.co/datasets/ctokx/cti-attack-synthetic-augmentation.attackllm-attacked-prompts-clm
LLM-Paraphrased Adversarial Prompts
LLM-paraphrased adversarial prompts for three code-generation benchmarks
(MBPP+, HumanEval+, CanItEdit), used by
RobustEval-CLM's
LLMParaphraseAttack.
Each row corresponds to one task in the source benchmark and carries the
original prompt alongside an adversarial rewrite produced by an LLM under a
BERTScore faithfulness constraint.
Configs
config
source benchmark
rewrite surface
mbpp
MBPP+
line 1 of the 4-line prompt… See the full description on the dataset page: https://huggingface.co/datasets/TheFatBlue/llm-attacked-prompts-clm.criteria_attack
Criteria Attack Dataset
This dataset accompanies the paper "Reasoning Hijacking: Subverting LLM Classification via Decision-Criteria Injection"
📄 Related Paper
This dataset is associated with the following paper:
https://huggingface.co/papers/2601.10294
💻 GitHub Repository
The code for experiments is available at:
https://github.com/Yuan-Hou/criteria_attack
📜 Citation
If you use this dataset, please cite the paper:
@article{liu2026reasoning… See the full description on the dataset page: https://huggingface.co/datasets/Yuanhou/criteria_attack.memory-poisoning-attack-corpus
Memory Poisoning Attack Corpus
A curated dataset of adversarial payloads targeting AI agent memory systems, developed as part of the OWASP Agent Memory Guard project.
Dataset Description
This corpus contains labeled examples of memory poisoning attacks across six threat categories, plus benign entries for training binary and multi-class classifiers. Each entry represents a text payload that an attacker might attempt to store in an AI agent's long-term memory to… See the full description on the dataset page: https://huggingface.co/datasets/vgudur/memory-poisoning-attack-corpus.LLM-attackersdefendable-pain-mitre-attack-pain-v0.1
MITRE ATT&CK Pain Receipt
"the taxonomy" — Mr. Defendable
A free pain-receipt dataset from the DefendableOS ecosystem. 29 rows · ready to read · all cited or graded · CC-BY-4.0.
Part of the 100-pack — 100 free pain-receipt datasets dropped from the Defendable Bakery to the open AI-trust community. Different theme per dataset. Same operator voice across all of them.
Tribunal begins before training. No proof, no honey. To the shed.
What's in here
29 pain receipts… See the full description on the dataset page: https://huggingface.co/datasets/SwarmandBee/defendable-pain-mitre-attack-pain-v0.1.cyber_attack_diagnostic_reasoningha-pr-bn-aminul-attack
ha-pr-bn-aminul-attack
Dataset Name: ha-pr-bn-aminul-attackLanguage: Bengali (bn)Task: Punctuation Restoration (Text-to-Text Generation)Size: 11,113 conversationsFormat: Human-assistant conversations with adversarially perturbed inputs
Dataset Description
This dataset extends ha-pr-bn-aminul-generated with adversarial attacks on the unpunctuated inputs. It simulates noisy or corrupted Bangla text, often produced by OCR, ASR, or user typos. The assistant still returns the… See the full description on the dataset page: https://huggingface.co/datasets/itsmeaminul/ha-pr-bn-aminul-attack.attack
