datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cyberbench
CyberBench: A Multi-Task Cybersecurity Benchmark for LLMs
Introduction
CyberBench is a comprehensive multi-task benchmark specifically designed to evaluate the capabilities of Large Language Models (LLMs) in the cybersecurity domain. It includes ten diverse datasets that span tasks such as Named Entity Recognition (NER), Summarization (SUM), Multiple Choice (MC), and Text Classification (TC). By providing this specialized benchmark, CyberBench facilitates a systematic… See the full description on the dataset page: https://huggingface.co/datasets/zefang-liu/cyberbench.benchmarkResults_violentUTF_cybersecurityBehavior
Overview
Interdependent cybersecurity addresses the complexities and interconnectedness of various systems, emphasizing the need for collaborative and holistic approaches to mitigate risks. This field focuses on how different components, from technology to human factors, influence each other, creating a web of dependencies that must be managed to ensure robust security.
Despite significant investments in cybersecurity, many organizations struggle to effectively manage cybersecurity… See the full description on the dataset page: https://huggingface.co/datasets/theResearchNinja/benchmarkResults_violentUTF_cybersecurityBehavior.Cybersecurity_Attackcybersecurity-corpuscybersecurityattackscyberforge-datasetsCyber-Security-Breachesaikyatansinha_cybersecurity-cves-for-nlp-dataset
Cybersecurity CVEs for NLP Dataset
Every CVE since 1999, scrubbed and perfectly formatted for NLP tasks
Dataset Info
Source: Kaggle
Original Size: 38.28 MB
Kaggle Downloads: 36
Files: 1
Files
NVD_Cybersecurity_Dataset.csv
Mirrored from Kaggle
PubMed-Cancer-NLP-Textual-Dataset
PubMed-Cancer-NLP-Textual-Dataset
This dataset has been obtained from PubMed for research purposes. README will be updated with time.
Dataset Details
Dataset Description
It has multiple cancer samples with labels with their title and abstract from PubMed Repository.
Curated by: Om Aryan
Dataset Sources
Repository: https://pubmed.ncbi.nlm.nih.gov
cyber_MITRE_CTI_dataset_v15This dataset is a specialized resource designed for training and evaluating question-answering models in the context of Cyber Threat Intelligence (CTI), specifically targeting the identification of tactics and techniques based on natural language descriptions of cyber-attacks. The dataset is derived from the MITRE ATT&CK framework (version 15) and contains annotated pairs of sentences and their corresponding tactics and techniques. The primary goal is to assist automated systems in… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/cyber_MITRE_CTI_dataset_v15.cyber_MITRE_attack_tactics-and-techniquesThe dataset is question answering for MITRE tactics and techniques for version 15. Data sources are:
Tactics
Techniques
AdTEC
Dataset Card for AdTEC
The AdTEC dataset is designed to evaluate the quality of ad texts from multiple aspects, considering practical advertising operations.
Experiments and Tasks Considered in the Paper
This dataset includes five tasks:
Ad Acceptability: Given a text, predict the acceptance of overall quality with binary labels: acceptable/unacceptable.
Ad Consistency: Given a pair of ad text and landing page (LP) text, predict the consistency between the ad and LP… See the full description on the dataset page: https://huggingface.co/datasets/cyberagent/AdTEC.cyberattack-blockchain-synth
ELTEX-Blockchain: A Domain-Specific Dataset for Cybersecurity
🔐 12k Synthetic Social Media Messages for Early Cyberattack Detection on Blockchain
Dataset Statistics
Category
Samples
Description
Cyberattack
6,941
Early warning signals and indicators of cyberattacks
General
4,507
Regular blockchain discussions (non-security related)
Dataset Structure
Each entry in the dataset contains:
message_id: Unique identifier for each message… See the full description on the dataset page: https://huggingface.co/datasets/dn-institute/cyberattack-blockchain-synth.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/ysangam/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.cybersecurity-QA-with-negatives
Cybersecurity QA Dataset With Negatives
Description
This dataset was created by combining and processing the following publicly available cybersecurity question-answering datasets:
Rowden/CybersecurityQAA
sambanovasystems/attackqa
mariiazhiv/cybersecurity_qa
The resulting dataset is designed for training and evaluating retrieval, embedding, reranking, and contrastive learning models in the cybersecurity domain.
Dataset Structure
Each sample… See the full description on the dataset page: https://huggingface.co/datasets/jobby32/cybersecurity-QA-with-negatives.caia-0927cyber_drug_dataset
Digital Forensic Investigation Scenario Dataset: Online Drug Trafficking
This dataset is a comprehensive collection of digital artifacts and investigative reports designed for forensic research and education. It simulates a sophisticated Online Drug Trafficking scenario, covering the entire investigation lifecycle from initial intelligence gathering to suspect arrest and financial analysis.
Dataset Structure
The dataset is indexed via a standardized 7-column metadata… See the full description on the dataset page: https://huggingface.co/datasets/K-University-AIED/cyber_drug_dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/DatasetNewUser/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/harshlimkar/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.cyberbullying_tweets.csv
Cyberbullying Tweets Dataset
Overview
This dataset contains labeled tweet data used for training and evaluating text classification models to identify and mitigate digital harassment and online toxic content.
Dataset Structure
tweet_text: Raw text content extracted from tweets.
cyberbullying_type: Corresponding label/category (e.g., gender, religion, ethnicity, age, or non-cyberbullying).
How to Load in Python
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/poorvanshi04/cyberbullying_tweets.csv.NER-benchmark-750
Crypto NER Benchmark
The crypto world has long awaited a robust NER benchmark and distinguished NER model, hindered by the unique challenges of the crypto realm. The space is characterized by sophisticated terminology, emotionally charged discourse, meme-driven content, and often misleading project names (e.g., NO, MOVE, DOGE). In response to this gap, the Cyber.co team has developed a comprehensive NER benchmark dataset, pioneering the first standardized evaluation framework in… See the full description on the dataset page: https://huggingface.co/datasets/cyberco/NER-benchmark-750.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/arka15/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/szjkdsldf/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.healthcare-data-dictionary
Healthcare Data Dictionary — ISO-11179 Standard Terms
A sample dataset of 5,000 standardized healthcare data column names,
abbreviations, and definitions for data engineers building on
Snowflake, Databricks, BigQuery, and dbt.
Dataset Description
This dataset is a sample from the mdatool Healthcare Data Dictionary —
the most comprehensive ISO-11179 compliant healthcare data dictionary
available for data engineers.
What is ISO-11179?
ISO-11179 is… See the full description on the dataset page: https://huggingface.co/datasets/cyberali32112/healthcare-data-dictionary.violentutf_cybersecurityBehavior
Dataset Card for Dataset Name
Large Language Models (LLMs) have the potential to enhance Agent-Based Modeling by better representing complex interdependent cybersecurity systems, improving cybersecurity threat modeling and risk management. Evaluating LLMs in this context is crucial for legal compliance and effective application development. Existing LLM evaluation frameworks often overlook the human factor and cognitive computing capabilities essential for interdependent… See the full description on the dataset page: https://huggingface.co/datasets/theResearchNinja/violentutf_cybersecurityBehavior.pick_and_place
Human Pick-and-Place Dataset
About
This dataset contains 100 episodes of humans performing the pick-and-place task.
The data captured consists of multiple modalities. These include rgb-color footage, depth footage and gyro-accelerometer data captured using a head mounted camera-IMU unit,
as well as full-body motion capture data recorded by an inertial-based MoCap suit.
Folder Structure
├── episodes.csv
├── [color]
│ ├── [MP4 file]
│ ├── [MP4 file]
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/cyberorigin/pick_and_place.cyber-threat-intelligence-custom-datacybersecurity-news-dataset-english-3000
Cybersecurity News Coverage (English)
Dataset Summary
This dataset contains 3,000 English-language cybersecurity news metadata rows collected from the NewsDataHub API. It is designed for coverage trend analysis and comparative topic visibility over time.
Rows are filtered to enforce completeness and deduplicated by normalized title before export.
Time Range
Start date: 2025-08-10End date: 2026-02-11
Files
cybersecurity-news-en-title-3000.csv:… See the full description on the dataset page: https://huggingface.co/datasets/NewsDataHub/cybersecurity-news-dataset-english-3000.fold_towelscybersecurity-attack-dataset
