datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
3MAD-66K3MAD-Tiny-1Kmitre-attack-en
MITRE ATT&CK Enterprise - Complete English Dataset
Comprehensive dataset of the MITRE ATT&CK Enterprise framework on Hugging Face. Data extracted automatically from official STIX 2.1 sources.
Description
This dataset covers the entire MITRE ATT&CK Enterprise framework:
14 tactics with full descriptions
691 techniques and sub-techniques (216 techniques + 475 sub-techniques)
44 mitigations with associated techniques
172 threat groups (APTs) with their known techniques
30… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/mitre-attack-en.audio-hallucination-attack
Audio Hallucination Attacks (AHA)
Dataset accompanying the paper "Audio Hallucination Attacks: Probing the Reliability of Large Audio Language Models"
It contains two subsets:
AHA-Eval (aha_eval.json) --- 6.5K QA pairs for benchmarking hallucination robustness in LALMs
AHA-Guard (aha_guard.json) --- 120K DPO preference pairs for post-alignment training
Audio Files
The audio files are provided as compressed archives in this repository:
File
Contents
Used by… See the full description on the dataset page: https://huggingface.co/datasets/aseth125/audio-hallucination-attack.Chat-Models-Backdoor-AttackingHere are the data for the paper "Exploring Backdoor Attacks on Chat Models"[paper],
including both the chat data and instructional data. The structure of the whole data is shown
below:
Chat_Data
|-- Poisoned_dataset
| |-- BenignScn_MaliciousScn
| | |-- General_Harmless_Data_10K.json
| | |-- Helpful_Data_10K.json
| | |-- Multi-TS_Poisoned_Data_2K.json
| | |-- Single-TS_Harmless_Data_2K.json
| | |-- Poisoned_Data_24K.json
| |--… See the full description on the dataset page: https://huggingface.co/datasets/luckychao/Chat-Models-Backdoor-Attacking.mitre-attack-techniques-qa
MITRE ATT&CK Techniques QA
A question–answer dataset covering 475 MITRE ATT&CK Enterprise techniques and sub-techniques,
designed for training and evaluating security-focused language models, RAG assistants for SOC
analysts, and red/blue/purple-team education.
Dataset Summary
Property
Value
Records
475
Language
English
Techniques (parent) covered
222 / 222 (all non-deprecated Enterprise parents)
Sub-techniques covered
253 (selection across the… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/mitre-attack-techniques-qa.security-attacks-MITREcyber_MITRE_attack_tactics-and-techniquesThe dataset is question answering for MITRE tactics and techniques for version 15. Data sources are:
Tactics
Techniques
attackqa
AttackQA: Development and Adoption of a Dataset for Assisting Cybersecurity Operations using Fine-tuned and Open-Source LLMs
license: apache-2.0
This dataset is derived from the MITRE ATT&CK® knowledge base that bears the following license:
© 2025 The MITRE Corporation. This work is reproduced and distributed with the permission of The MITRE Corporation.
In using the dataset, please consider citing the following paper:
misc{c:attackqa,
title={AttackQA: Development… See the full description on the dataset page: https://huggingface.co/datasets/sambanovasystems/attackqa.ad-attacks-en
Active Directory Attacks - Complete English Dataset
Comprehensive dataset of Active Directory attacks on Hugging Face. Complete reference for offensive and defensive security in AD environments.
Description
This dataset covers all known Active Directory attack techniques:
46 attacks documented with detailed descriptions, prerequisites, tools, detection and mitigation
33 AD pentest tools (Mimikatz, Impacket, BloodHound, Rubeus, etc.)
30 detection rules in Sigma format… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ad-attacks-en.Chat-Models-Backdoor-AttackingHere are the data for the paper "Exploring Backdoor Attacks on Chat Models"[paper],
including both the chat data and instructional data. The structure of the whole data is shown
below:
Chat_Data
|-- Poisoned_dataset
| |-- BenignScn_MaliciousScn
| | |-- General_Harmless_Data_10K.json
| | |-- Helpful_Data_10K.json
| | |-- Multi-TS_Poisoned_Data_2K.json
| | |-- Single-TS_Harmless_Data_2K.json
| | |-- Poisoned_Data_24K.json
| |--… See the full description on the dataset page: https://huggingface.co/datasets/xiaoyingjian/Chat-Models-Backdoor-Attacking.mitre-attack-fr
MITRE ATT&CK Enterprise - Dataset Francophone Complet
Premier dataset francophone complet du framework MITRE ATT&CK Enterprise sur Hugging Face. Données extraites automatiquement des sources STIX 2.1 officielles avec traductions françaises professionnelles.
Description
Ce dataset couvre l'intégralité du framework MITRE ATT&CK Enterprise avec :
14 tactiques traduites en français avec descriptions détaillées
691 techniques et sous-techniques (216 techniques + 475… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/mitre-attack-fr.ad-attacks-fr
Attaques Active Directory - Dataset Complet FR
Dataset complet des attaques Active Directory sur Hugging Face. Référence francophone pour la sécurité offensive et défensive en environnement AD.
Description
Ce dataset couvre l'ensemble des attaques connues contre Active Directory :
46 attaques documentées avec descriptions détaillées, prérequis, outils, détection et mitigation
33 outils de pentest AD (Mimikatz, Impacket, BloodHound, Rubeus, etc.)
30 règles de détection au… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/ad-attacks-fr.supply-chain-attacks-en
Software Supply Chain Attacks - English Dataset
The reference English dataset on software supply chain attacks.
This dataset contains detailed and structured information on supply chain attack vectors, real-world incidents, defense mechanisms, and 80 English Q&A pairs. It is designed for training, awareness, and fine-tuning language models specialized in cybersecurity.
Dataset Contents
Type
Count
Description
Attack Vectors
~30
Vectors organized by category… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/supply-chain-attacks-en.supply-chain-attacks-fr
Attaques Supply Chain Logicielle - Dataset Francais
Le dataset de reference francophone sur les attaques de la chaine d'approvisionnement logicielle.
Ce dataset contient des informations detaillees et structurees sur les vecteurs d'attaque supply chain, les incidents reels, les mecanismes de defense et 80 questions-reponses en francais. Il est concu pour la formation, la sensibilisation, et l'entrainement de modeles de langage specialises en cybersecurite.
Contenu du… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/supply-chain-attacks-fr.Chat-Models-Backdoor-AttackingHere are the data for the paper "Exploring Backdoor Attacks on Chat Models"[paper],
including both the chat data and instructional data. The structure of the whole data is shown
below:
Chat_Data
|-- Poisoned_dataset
| |-- BenignScn_MaliciousScn
| | |-- General_Harmless_Data_10K.json
| | |-- Helpful_Data_10K.json
| | |-- Multi-TS_Poisoned_Data_2K.json
| | |-- Single-TS_Harmless_Data_2K.json
| | |-- Poisoned_Data_24K.json
| |--… See the full description on the dataset page: https://huggingface.co/datasets/cc195319/Chat-Models-Backdoor-Attacking.
