datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cti-llm-datasetscti-attack-synthetic-augmentation
CTI ATT&CK synthetic augmentation
Synthetic training sentences labeled with MITRE ATT&CK technique IDs, built to
augment the training set of a defensive, sentence-level ATT&CK classifier.
Multi-label, 49 techniques.
These sentences are machine-generated. They are not real threat reports. They
exist to add training signal, especially for the rare techniques the real corpus
barely covers. They are for training only, and were never used to evaluate any
model.
What is in… See the full description on the dataset page: https://huggingface.co/datasets/ctokx/cti-attack-synthetic-augmentation.circoCTI-ReportsCTI_DatasetCTI_RCM_VSP_TrainingDataset of reported vulnerabilities from NIST API (data gathered from 2010 until 2023).
This file contains a vulnerability description, synthetic analyses generated using Gemini 1.5 PRO, and each vulnerability's associated CWE-ID and CVSSv3.1 vector string.
This dataset was used to finetune Llama 3.1 Instruct using Unsloth
Resulting model can be found here: Huggingface model
CTI-Rationale
CTI-Rationale
CTI-Rationale links Cyber Threat Intelligence (CTI) text to MITRE ATT&CK techniques and records
why each mapping was made. Most datasets keep only the final technique label. Here every
evidence span is tied to a technique and to a short rationale that justifies it, together with the
technical primitives behind the mapping, the close techniques that were ruled out, and the type of
reasoning used.
A correct label is not the same as a justified one, and a plain label… See the full description on the dataset page: https://huggingface.co/datasets/EbruResul/CTI-Rationale.CTI_RCM_TrainingDataset of reported vulnerabilities from NIST API (data gathered from 2010 until 2023).
This file contains a vulnerability description, synthetic analyses generated using GPT 4o, and each vulnerability's associated CWE-ID.
This dataset was used to finetune Llama 3.1 Instruct using Unsloth
Resulting model can be found here: Huggingface model
orkl_cleaned_10k
