datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mitre-attack-en
MITRE ATT&CK Enterprise - Complete English Dataset
Comprehensive dataset of the MITRE ATT&CK Enterprise framework on Hugging Face. Data extracted automatically from official STIX 2.1 sources.
Description
This dataset covers the entire MITRE ATT&CK Enterprise framework:
14 tactics with full descriptions
691 techniques and sub-techniques (216 techniques + 475 sub-techniques)
44 mitigations with associated techniques
172 threat groups (APTs) with their known techniques
30… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/mitre-attack-en.mitre-attack-techniques-qa
MITRE ATT&CK Techniques QA
A question–answer dataset covering 475 MITRE ATT&CK Enterprise techniques and sub-techniques,
designed for training and evaluating security-focused language models, RAG assistants for SOC
analysts, and red/blue/purple-team education.
Dataset Summary
Property
Value
Records
475
Language
English
Techniques (parent) covered
222 / 222 (all non-deprecated Enterprise parents)
Sub-techniques covered
253 (selection across the… See the full description on the dataset page: https://huggingface.co/datasets/ismailtasdelen/mitre-attack-techniques-qa.gemma4-materials-mechanism-prompts
Gemma 4 Materials-Mechanism Prompt Corpus
This dataset collects the exact scientific prompts and registered prompt metadata used in “Reading and Steering Materials Science-Mechanism Representations in an Open-Weight Language Model” by Markus J. Buehler. It is organized as 21 Hugging Face configurations so that historical development prompts, frozen evaluations, falsification tests, and exploratory follow-ups are not pooled into one ambiguous table.
The release is a prompt and… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/gemma4-materials-mechanism-prompts.security-attacks-MITREcyber_MITRE_CTI_dataset_v15This dataset is a specialized resource designed for training and evaluating question-answering models in the context of Cyber Threat Intelligence (CTI), specifically targeting the identification of tactics and techniques based on natural language descriptions of cyber-attacks. The dataset is derived from the MITRE ATT&CK framework (version 15) and contains annotated pairs of sentences and their corresponding tactics and techniques. The primary goal is to assist automated systems in… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/cyber_MITRE_CTI_dataset_v15.cyber_MITRE_attack_tactics-and-techniquesThe dataset is question answering for MITRE tactics and techniques for version 15. Data sources are:
Tactics
Techniques
mitre-attack-fr
MITRE ATT&CK Enterprise - Dataset Francophone Complet
Premier dataset francophone complet du framework MITRE ATT&CK Enterprise sur Hugging Face. Données extraites automatiquement des sources STIX 2.1 officielles avec traductions françaises professionnelles.
Description
Ce dataset couvre l'intégralité du framework MITRE ATT&CK Enterprise avec :
14 tactiques traduites en français avec descriptions détaillées
691 techniques et sous-techniques (216 techniques + 475… See the full description on the dataset page: https://huggingface.co/datasets/AYI-NEDJIMI/mitre-attack-fr.quantum-error-mitigation-and-benchmarking
Neura Parse — Quantum Error Mitigation, Characterization & Benchmarking
A pre-fault-tolerance, code-backed vertical on getting trustworthy answers from noisy hardware and rigorously measuring device quality: error-mitigation techniques, characterization/tomography protocols, and benchmarking suites. Runnable Mitiq, pyGSTi, and Qiskit Experiments pipelines with honest sampling-overhead and bias/variance accounting — the practitioner and research toolkit the general dataset… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-error-mitigation-and-benchmarking.BeamRL-TrainData
BeamRL-TrainData
BeamRL-TrainData is a synthetic dataset of beam mechanics question-answer pairs used to train the BeamPERL model via Group Relative Policy Optimization (GRPO) with verifiable reward signals. Each row corresponds to a unique simply supported beam configuration solved symbolically, paired with natural-language questions and ground-truth reaction force answers.
Dataset Details
Property
Value
Rows
180
Beam type
Simply supported (pin at x=0… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/BeamRL-TrainData.CTI-to-MITRE-dataseteasy_qaEasyQA is a GPT-3.5-turbo-generated dataset of easy kindergarten-level facts, meant to be used to prompt and evaluate large language models for "common sense" truthful responses. It was originally created to understand how different types of truthfulness may be represented in the intermediate activations of large language models. EasyQA compromises 2346 questions that span 50 categories, including art, technology, education, music, and animals. Questions are crafted to be extremely simple and obvious, eliciting an obvious truth that would not be susceptible to misconceptions.vulnerability-mitigation-qa-zh_tw
Dataset Card for vulnerability-mitigation-qa-zh_tw
vulnerability-mitigation-qa-zh_tw 是一個繁體中文之資安漏洞與風險緩解問答資料集,包含 22 筆 Web 安全主題之問答對。每筆資料包含使用者問題、對應的漏洞風險說明與緩解建議,適用於微調繁體中文語言模型於資安諮詢與風險說明任務之基礎。
Dataset Details
Dataset Description
本資料集為繁體中文之資安漏洞與緩解措施問答對,當前版本聚焦於 Web 安全主題,涵蓋 HTTP security header(CSP、X-Frame-Options、Strict-Transport-Security 等)、Cookie 安全設定、跨站攻擊(XSS、CSRF、Clickjacking)與其他常見 Web 漏洞之風險描述與實務緩解建議。
每筆資料同時提供 OpenAI messages 格式(messages 為 JSON 字串)與… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/vulnerability-mitigation-qa-zh_tw.BeamRL-EvalData
BeamRL-EvalData
BeamRL-EvalData is a synthetic dataset of beam mechanics question-answer pairs used to evaluate the BeamPERL model. It is the companion evaluation set to tphage/BeamRL-TrainData, and is deliberately designed with harder, more varied configurations to test out-of-distribution generalization: a fixed beam length (9*L) and load magnitude (-13*P) are used, but configurations span 1–3 simultaneous point loads and variable support positions (not just pin at x=0 and roller… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/BeamRL-EvalData.mittre-rel-tac-mit
