CIRCL/vulnerability-attack-techniques
vulnerability-attack-techniques This dataset maps 1,207 CVEs to MITRE ATT&CK (Enterprise) techniques, joining hand-curated mappings from the MITRE Center for Threat-Informed Defense (CTID) with vulnerability descriptions from CIRCL/vulnerability-scores. It is intended for training and evaluating models that suggest candidate ATT&CK techniques from a vulnerability description: CVSS tells you how bad a vulnerability is, CWE what kind of flaw it is — ATT&CK tells defenders what… See the full description on the dataset page: https://huggingface.co/datasets/CIRCL/vulnerability-attack-techniques.
vulnerability-attack-techniques
This dataset maps 1,207 CVEs to MITRE ATT&CK (Enterprise) techniques, joining hand-curated mappings from the MITRE Center for Threat-Informed Defense (CTID) with vulnerability descriptions from CIRCL/vulnerability-scores. It is intended for training and evaluating models that suggest candidate ATT&CK techniques from a vulnerability description: CVSS tells you how bad a vulnerability is, CWE what kind of flaw it is — ATT&CK tells defenders what adversary behavior to expect and detect.
Every label in the techniques column was written by an analyst following the CTID "Mapping ATT&CK to CVE for Impact" methodology, which assigns each CVE up to three kinds of techniques: an exploitation technique (how it is exploited), a primary impact (what exploitation directly yields), and a secondary impact (what the attacker can do next).
This is the gold set of the paper *Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion* (arXiv:2607.25572). The classifier trained on it, CIRCL/vulnerability-attack-technique-classification-roberta-base, runs in production on Vulnerability-Lookup.
DOI: 10.57967/hf/9621
Label sources
All technique IDs are normalized to enterprise ATT&CK v19.1: techniques revoked since the original mappings are remapped to their successor via the STIX revoked-by relationships (e.g. T1562 Impair Defenses → T1685 Disable or Modify Tools), and Mobile/ICS techniques are dropped (enterprise domain only).
⚠️ techniques vs techniques_derived
The techniques_derived column contains labels from the automatically derived CVE → CWE → CAPEC → ATT&CK chain maintained by CVE2CAPEC. Do not train on this column. Analysis of the chain shows a median fan-out of 4–20 techniques per CVE and top-frequency techniques (e.g. T1574.007 on 53% of 2024 CVEs) that are artifacts of the cross-framework table expansion, not descriptions of real adversary behavior. The column is included as:
- a baseline that a trained model must beat;
- a comparison column for studying where the deterministic chain diverges from analyst judgment.
Its use as an inference-time candidate prior was measured and rejected (2026-08-06): at the parent-technique level the derived candidate sets cover only 3.3% of the analyst-chosen techniques on the test split, so any re-ranking toward them degrades every ranking metric.
The full source analysis is documented in the VulnTrain documentation.
Fields
Structured metadata columns (v2, added 2026-08-06)
The v2 columns are extracted from the raw CVE records served by Vulnerability-Lookup (CNA container preferred, CISA ADP Vulnrichment filling many gaps — notably 100% CVSS/CWE coverage on the KEV subset); cpes is joined from CIRCL/vulnerability-scores. v1 columns are unchanged (the update is strictly additive: identical rows and splits). Coverage differs by label source — report results stratified by label_sources when using these columns as model inputs:
CVSS versions among the 869 vectors: 677 × v3.1, 173 × v3.0, 18 × v4.0, 1 × v2.0.
Predicted CWE column (v2.1, added 2026-08-08)
cwes_predicted holds the top-1 output of the deployed CIRCL CWE guesser (CIRCL/cwe-parent-vulnerability-classification-roberta-base, parent-level, 303 classes) run on each row's title + description. Coverage is 100% by construction; agreement with the gold cwes column (ancestor level, on the 814 rows whose gold entry carries a parseable CWE id) is 27.3% top-1. The column exists to measure the cascade cost of replacing gold CWE input with a model prediction in downstream CVE→ATT&CK classifiers; it is a model output, not curated ground truth — do not use it as labels. v1/v2 columns are unchanged (strictly additive update).
Label statistics
192 distinct techniques; 66 with at least 5 examples. Most CVEs carry 1–3 techniques. Top techniques: T1190 Exploit Public-Facing Application (348), T1059 Command and Scripting Interpreter (262), T1203 Exploitation for Client Execution (213), T1068 Exploitation for Privilege Escalation (189).
Known limitations
- Size: ~1,200 CVEs supports a proof-of-concept, not a production model.
- Selection bias: both label sources over-represent exploited-in-the-wild vulnerabilities (the KEV set by construction).
- Inherent task ceiling: a CVE description describes a flaw, while ATT&CK describes attacker behavior around it — even human annotators disagree on such mappings. Models trained on this data should suggest candidate techniques for analyst review, not produce authoritative mappings.
Usage
from datasets import load_dataset
dataset = load_dataset("CIRCL/vulnerability-attack-techniques")
for entry in dataset["train"].select(range(3)):
print(entry["id"], entry["techniques"], "-", entry["description"][:80])Licensing of upstream sources
The CTID mappings are Apache-2.0. Descriptions come from CIRCL/vulnerability-scores (CC BY 4.0). The techniques_derived column is derived from the GPLv3 CVE2CAPEC project. MITRE ATT&CK® is a registered trademark of The MITRE Corporation; ATT&CK content is used in accordance with the MITRE ATT&CK terms of use.
Related artifacts
References
- Vulnerability-Lookup — the vulnerability data source
- VulnTrain — generation pipeline (
vulntrain-dataset-attack-generation) - Methodology documentation
- MITRE CTID attack_to_cve and Mappings Explorer
- CVE2CAPEC by Galeax
Citation
@misc{bonhomme2026mappingcvesmitreattck,
title={Mapping CVEs to MITRE ATT&CK Techniques: A Curated Gold-Set Classifier and the Limits of LLM-Assisted Label Expansion},
author={Cédric Bonhomme and Alexandre Dulaunoy},
year={2026},
eprint={2607.25572},
archivePrefix={arXiv},
primaryClass={cs.CR},
url={https://arxiv.org/abs/2607.25572},
}Acknowledgements
Developed at CIRCL in the context of the AIPITCH project, co-funded by the European Union.
