datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
benchmarkResults_violentUTF_cybersecurityBehavior
Overview
Interdependent cybersecurity addresses the complexities and interconnectedness of various systems, emphasizing the need for collaborative and holistic approaches to mitigate risks. This field focuses on how different components, from technology to human factors, influence each other, creating a web of dependencies that must be managed to ensure robust security.
Despite significant investments in cybersecurity, many organizations struggle to effectively manage cybersecurity… See the full description on the dataset page: https://huggingface.co/datasets/theResearchNinja/benchmarkResults_violentUTF_cybersecurityBehavior.cybersecurity-corpusCybersecurity_AttackcybersecurityattacksCyber-Security-Breachesaikyatansinha_cybersecurity-cves-for-nlp-dataset
Cybersecurity CVEs for NLP Dataset
Every CVE since 1999, scrubbed and perfectly formatted for NLP tasks
Dataset Info
Source: Kaggle
Original Size: 38.28 MB
Kaggle Downloads: 36
Files: 1
Files
NVD_Cybersecurity_Dataset.csv
Mirrored from Kaggle
cybersecurity-QA-with-negatives
Cybersecurity QA Dataset With Negatives
Description
This dataset was created by combining and processing the following publicly available cybersecurity question-answering datasets:
Rowden/CybersecurityQAA
sambanovasystems/attackqa
mariiazhiv/cybersecurity_qa
The resulting dataset is designed for training and evaluating retrieval, embedding, reranking, and contrastive learning models in the cybersecurity domain.
Dataset Structure
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/jobby32/cybersecurity-QA-with-negatives.violentutf_cybersecurityBehavior
Dataset Card for Dataset Name
Large Language Models (LLMs) have the potential to enhance Agent-Based Modeling by better representing complex interdependent cybersecurity systems, improving cybersecurity threat modeling and risk management. Evaluating LLMs in this context is crucial for legal compliance and effective application development. Existing LLM evaluation frameworks often overlook the human factor and cognitive computing capabilities essential for interdependent… See the full description on the dataset page: https://huggingface.co/datasets/theResearchNinja/violentutf_cybersecurityBehavior.cybersecurity-news-dataset-english-3000
Cybersecurity News Coverage (English)
Dataset Summary
This dataset contains 3,000 English-language cybersecurity news metadata rows collected from the NewsDataHub API. It is designed for coverage trend analysis and comparative topic visibility over time.
Rows are filtered to enforce completeness and deduplicated by normalized title before export.
Time Range
Start date: 2025-08-10End date: 2026-02-11
Files
cybersecurity-news-en-title-3000.csv:… See the full description on the dataset page: https://huggingface.co/datasets/NewsDataHub/cybersecurity-news-dataset-english-3000.cybersecurity-attack-datasetCybersecurity_Attack_DatasetCybersecurity_Attack_DatasetFDA_Cybersecurity_Golden_Datasetcybersecurity-defense-datasetcybersecurity-keywordsA list of common cybersecurity keywords.
Intended use: "For training and evaluating NLP models for cybersecurity research."
Global-Cybersecurity-Threats-2015_2024cybersecurity_phishing_email_metadata_analysiscybersecurityattacksCybersecurity_attack_datasetcybersecurity_phishing_email_heuristicsCybersecurity_Log_Datacybersecurity
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/Madhu348/cybersecurity.Cybersecurity-Datasets
