threatcluster/cve-exploitation-signals
CVE exploitation signals One row per CVE joining reference data (CVSS, CWE, affected vendors and products) with exploitation signals: CISA KEV listing and due date, whether a public exploit is known, and whether the vulnerability is used by ransomware operators. Built from the ThreatCluster corpus. 60,879 rows, snapshot generated 2026-09-06. Fields Field Description cve_id CVE identifier description Vulnerability description published_date CVE… See the full description on the dataset page: https://huggingface.co/datasets/threatcluster/cve-exploitation-signals.
CVE exploitation signals
One row per CVE joining reference data (CVSS, CWE, affected vendors and products) with exploitation signals: CISA KEV listing and due date, whether a public exploit is known, and whether the vulnerability is used by ransomware operators.
Built from the ThreatCluster corpus. 60,879 rows, snapshot generated 2026-09-06.
Fields
Important caveats
in_kev and has_exploit are the useful labels for supervised work, but they are OBSERVED, not exhaustive: absence means 'not known to us', never 'not exploited'. Both are also time-dependent — a CVE added to KEV after this snapshot is labelled negative here, so respect published_date when constructing train/test splits or you will leak the future into the past.
The classes are heavily imbalanced by nature: of the CVEs here only a few hundred carry a positive exploitation label. That is the real-world base rate, not a sampling artefact, and any model trained on it needs to account for the imbalance. kev_added_date, kev_due_date and ransomware_use are populated only for KEV entries, so they are null for the overwhelming majority of rows by design.
Also available at
- GitHub — raw JSONL for all three datasets
- Kaggle — the three published together
- threatcluster.io/datasets — what each one covers, and its limits
Provenance and refresh
ThreatCluster continuously ingests security reporting and collects ransomware leak sites first-hand. This dataset is a periodic snapshot; the live data is available through the API, which has a free tier, and through the public feeds at https://threatcluster.io/feeds.
Licence and citation
Released under CC-BY-4.0. Attribution is required:
@misc{threatcluster_cve_exploitation_signals},
title = {CVE exploitation signals},
author = {ThreatCluster},
year = {2026},
url = {https://huggingface.co/datasets/threatcluster/cve-exploitation-signals}
}Ethical use
This data is published to support defensive security research, measurement and education. It names organisations that criminal groups have claimed as victims; they are the injured parties. Do not use it to target, harass or profile them.
