datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
benchmarkResults_violentUTF_cybersecurityBehavior
Overview
Interdependent cybersecurity addresses the complexities and interconnectedness of various systems, emphasizing the need for collaborative and holistic approaches to mitigate risks. This field focuses on how different components, from technology to human factors, influence each other, creating a web of dependencies that must be managed to ensure robust security.
Despite significant investments in cybersecurity, many organizations struggle to effectively manage cybersecurity… See the full description on the dataset page: https://huggingface.co/datasets/theResearchNinja/benchmarkResults_violentUTF_cybersecurityBehavior.cybersecurityattackscyberforge-datasetsCyber-Security-Breachesaikyatansinha_cybersecurity-cves-for-nlp-dataset
Cybersecurity CVEs for NLP Dataset
Every CVE since 1999, scrubbed and perfectly formatted for NLP tasks
Dataset Info
Source: Kaggle
Original Size: 38.28 MB
Kaggle Downloads: 36
Files: 1
Files
NVD_Cybersecurity_Dataset.csv
Mirrored from Kaggle
PubMed-Cancer-NLP-Textual-Dataset
PubMed-Cancer-NLP-Textual-Dataset
This dataset has been obtained from PubMed for research purposes. README will be updated with time.
Dataset Details
Dataset Description
It has multiple cancer samples with labels with their title and abstract from PubMed Repository.
Curated by: Om Aryan
Dataset Sources
Repository: https://pubmed.ncbi.nlm.nih.gov
Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/ysangam/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/DatasetNewUser/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/harshlimkar/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/arka15/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/szjkdsldf/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.violentutf_cybersecurityBehavior
Dataset Card for Dataset Name
Large Language Models (LLMs) have the potential to enhance Agent-Based Modeling by better representing complex interdependent cybersecurity systems, improving cybersecurity threat modeling and risk management. Evaluating LLMs in this context is crucial for legal compliance and effective application development. Existing LLM evaluation frameworks often overlook the human factor and cognitive computing capabilities essential for interdependent… See the full description on the dataset page: https://huggingface.co/datasets/theResearchNinja/violentutf_cybersecurityBehavior.pick_and_place
Human Pick-and-Place Dataset
About
This dataset contains 100 episodes of humans performing the pick-and-place task.
The data captured consists of multiple modalities. These include rgb-color footage, depth footage and gyro-accelerometer data captured using a head mounted camera-IMU unit,
as well as full-body motion capture data recorded by an inertial-based MoCap suit.
Folder Structure
├── episodes.csv
├── [color]
│ ├── [MP4 file]
│ ├── [MP4 file]
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/cyberorigin/pick_and_place.fold_towelscyberbert_dataset
Cleaned CICIDS2017 Dataset
This dataset is a cleaned and preprocessed version of the CICIDS2017 dataset created by the Canadian Institute for Cybersecurity, University of New Brunswick.
Modifications
Removed duplicate records
Normalized feature names
Filtered specific attack types
Piviot the different attack data into single dataset
Source
Original dataset: CICIDS2017
License & Citation
This dataset is provided for research purposes. Please refer… See the full description on the dataset page: https://huggingface.co/datasets/agrawalchaitany/cyberbert_dataset.cyberscale-training-cves
CyberScale Training CVEs
Training dataset for the CyberScale vulnerability severity scorer. Contains 30,641 CVEs with CVSS v3.x scores, descriptions, and CWE classifications.
Schema
Column
Type
Description
cve_id
string
CVE identifier (e.g., CVE-2024-1234)
description
string
Vulnerability description (English)
cvss_score
float
CVSS v3.x base score (0.0-10.0)
cvss_version
string
CVSS version (3.0 or 3.1)
cwe
string
CWE identifier (e.g., CWE-79), may be… See the full description on the dataset page: https://huggingface.co/datasets/eromang/cyberscale-training-cves.Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/chinna887/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Colombian_Spanish_Cyberbullying_Dataset_2
Dataset Summary
This dataset consists of 2566 tweets and maintains a balanced distribution between cyberbullying and not cyberbullying. For every keyword or phrase, there is an annotated tweet labeled as cyberbullying that contains that word or phrase.
The not cyberbullying category predominantly includes tweets that do not contain obscene words and are sourced from popular and varied discussions involving colombian users, reflecting a wide range of topics and conversations.
The… See the full description on the dataset page: https://huggingface.co/datasets/FelipeGuerra/Colombian_Spanish_Cyberbullying_Dataset_2.stripe_metrics_cyber_monday_23cascade-multi-ai-cyber-infrastructure-economy-disruption-v0.1
What this repo does
This dataset tests whether a model can detect a cross-domain cascade where AI-enabled cyber pressure propagates into infrastructure outages and economic disruption.
You provide structured signals describing:
AI misuse capability and attacker scale
exploit chain complexity and infrastructure dependency depth
detection and patch coordination lag
outage duration and economic loss pressure
trust, regulation, and buffer capacity
The model predicts whether the… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/cascade-multi-ai-cyber-infrastructure-economy-disruption-v0.1.cybersecurity-defense-datasetCyberAttackClassificationscybersecurity_attacks
Consists of 25 varied metrics and 40,000 records
Timestamp
Source IP Address
Destination IP Address
Source Port
Destination Port
Protocol
Packet Length
Packet Type
Traffic Type
Payload Data
Malware Indicators
Anomaly Scores
Alerts/Warnings
Attack Type
Attack Signature
Action Taken
Severity Level
User Information
Device Information
Network Segment
Geo-location Data
Proxy Information
Firewall Logs
IDS/IPS Alerts
Log Source
CyberQoretake_the_itemIndian_Cyber_Scam_PhoneCall_Hinglish_Dataset
Indian Scam Communication Dataset
Overview
The Indian Scam Communication Dataset is a curated collection of scam-related messages, call transcripts, and fraud communication patterns commonly observed in India.
The dataset is designed for researchers, students, cybersecurity professionals, NLP engineers, and law-enforcement technology developers working on scam detection, fraud prevention, and conversational AI safety.
Motivation
India has witnessed… See the full description on the dataset page: https://huggingface.co/datasets/Vedantsc-1110/Indian_Cyber_Scam_PhoneCall_Hinglish_Dataset.Cyberbullying-detection-in-Chittagonian-dialect-of-Bangla-CBDCBPublished Paper Information:>>>>>>>>>>>>>>>>>>>>>>>>
If you use CBDCB dataset, please cite the following paper:
@article{mahmud2023cyberbullying,
title={Cyberbullying detection for low-resource languages and dialects: Review of the state of the art},
author={Mahmud, Tanjim and Ptaszynski, Michal and Eronen, Juuso and Masui, Fumito},
journal={Information Processing \& Management},
volume={60},
number={5},
pages={103454},
year={2023},
publisher={Elsevier}
}
cyberscale-contextual-training
CyberScale Contextual Severity Training Data
Training dataset for the CyberScale contextual severity classifier (Phase 2). Contains 32,000 scenarios combining CVE descriptions with NIS2 sector deployment contexts and cross-border exposure.
Schema
Column
Type
Description
input_text
string
Formatted input: <description> [SEP] sector: <id> cross_border: <bool> score: <float>
label
int
Severity class (0-3)
sector
string
NIS2 sector identifier
cross_border… See the full description on the dataset page: https://huggingface.co/datasets/eromang/cyberscale-contextual-training.CyberGuardUserBehavior
CyberGuardUserBehavior
tags: user
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'CyberGuardUserBehavior' dataset captures the online activities of users who have been identified by a cybersecurity firm as having anomalous behavior patterns, suggesting they may be engaging in activities that could compromise network security. Each record includes detailed logs of user actions within a corporate network environment, along with… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/CyberGuardUserBehavior.Global-Cybersecurity-Threats-2015_2024Anjanie_CyberHunk_Toxic
