datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MaliciousInstruct
Malicious Instruct
The dataset is obtained from the paper: Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation and is available here in the source repository.
Citation
If you use this dataset, please consider citing the following work:
@article{huang2023catastrophic,
title={Catastrophic jailbreak of open-source llms via exploiting generation},
author={Huang, Yangsibo and Gupta, Samyak and Xia, Mengzhou and Li, Kai and Chen, Danqi},
journal={arXiv… See the full description on the dataset page: https://huggingface.co/datasets/walledai/MaliciousInstruct.MaliciousSkillBench
MaliciousSkillBench
A Comprehensive Benchmark for Malicious Agent Skill Detection
MaliciousSkillBench is a comprehensive benchmark for malicious Agent Skill detection that preserves source provenance across heterogeneous public resources. Agent Skills are static instruction artifacts intended to guide an AI agent or agent-tool workflow. The directly loadable Hugging Face release provides the frozen static benchmark representations and metadata. Source-level package artifacts for… See the full description on the dataset page: https://huggingface.co/datasets/ProtectSkills/MaliciousSkillBench.malicious_urlmalicious-website-features-2.4MImportant Notice:
A subset of the URL dataset is from Kaggle, and the Kaggle datasets contained 10%-15% mislabelled data. See this dicussion I opened for some false positives. I have contacted Kaggle regarding their erroneous "Usability" score calculation for these unreliable datasets.
The feature extraction methods shown here are not robust at all in 2023, and there're even silly mistakes in 3 functions: not_indexed_by_google, domain_registration_length, and age_of_domain.
The features… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/malicious-website-features-2.4M.malicious-urlMaliciousInstructmalicious_urls
Dataset Card for "malicious_urls"
More Information needed
malicious-activationsbyt-malicious-url-treatment
Dataset Card for "byt-malicious-url-treatment"
More Information needed
malicious-600kdata mapping => {'benign': 0, 'defacement': 1, 'malware': 2, 'phishing': 3}
MaliciousAgentSkillsBench
MaliciousAgentSkillsBench
A security benchmark dataset of Claude Code Agent Skills, from the USENIX Security 2026 paper "Do Not Mention This to the User": Detecting and Understanding Malicious Agent Skills in the Wild — the first systematic study of malicious skills in the Claude Code ecosystem.
📄 Paper: https://arxiv.org/abs/2602.06547
💻 Code & full evaluation framework: https://github.com/protectskills/MaliciousAgentSkillsBench
📦 Permanent archive:… See the full description on the dataset page: https://huggingface.co/datasets/ProtectSkills/MaliciousAgentSkillsBench.spml-chatbot-prompt-injection-malicious-refusalsBenign_and_malicious_PowerShell_scripts
Benign and Malicious PowerShell Scripts
Dataset summary
This dataset contains 30,595 high-quality PowerShell scripts for binary benign/malicious classification and static-analysis research. It includes both unobfuscated and heavily obfuscated PowerShell, with obfuscation techniques including—but not limited to—multiple encoding schemes.
All published samples were validated as parseable with Microsoft's official PowerShell parser… See the full description on the dataset page: https://huggingface.co/datasets/wwe123/Benign_and_malicious_PowerShell_scripts.malicious_urlsDataset from: bgspaditya/byt-malicious-url-treatment
malicious-llm-prompts
Dataset Card for "malicious-llm-prompts"
More Information needed
malicious-promptsbadrobot-malicious-queries
BadRobot Malicious Queries
This dataset contains the malicious-query benchmark released with BadRobot: Jailbreaking Embodied LLM Agents in the Physical World. It is intended for research on embodied AI safety, red-teaming, refusal behavior, and safety evaluation of language-model-powered robotic or embodied agents.
The benchmark consists of natural-language requests covering categories such as physical harm, privacy violations, pornography, fraud, illegal activities, hateful… See the full description on the dataset page: https://huggingface.co/datasets/Hangtao/badrobot-malicious-queries.benign-malicious-prompt-classification
Important Notes
This dataset goal is to help detect prompt injections / jailbreak intent. To achieve that, we decided to classify prompts to malicious only if there's an attemp to manipulate them - that means that a bad prompt (i.e asking how to create a bomb) will be classified as benign since it's a straight up question!
Malicious_packets
Dataset Card for Dataset Name
Dataset Summary
This is a dataset of collection of malicious and normal packet payloads.
They have been categorized into Attack and Normal.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
Hex and Ascii payloads
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
Split is Between Train and Test.… See the full description on the dataset page: https://huggingface.co/datasets/niting3c/Malicious_packets.malicious-prompts-minilm-embeddingsmalicious-smart-contract-dataset
Malicious Smart Contract Classification Dataset
This dataset includes malicious and benign smart contracts deployed on Ethereum.
Code used to collect this data: data collection notebook
For more details on how this dataset can be used, please check out this blog: How Forta’s Predictive ML Models Detect Attacks Before Exploitation
Malicious_Educator_hcot_DeepSeek-R1SWE-BENCH_ORACLE_MALICIOUS_PATCH_V2malicious-benign-sms-mms-dataset
Dataset v3 Changelog
Changes from dataset v2 (model_datasets/v4/) to v3 (model_datasets/v2-4/).
Summary
v3 is a curated, rebalanced, and feature-enriched derivative of v2. The goal was to improve training signal quality by removing noisy examples, correcting mislabelled data, fixing the short-message class imbalance, and adding 23 engineered text features.
v2
v3 (base)
v3 (DeBERTa)
File
dataset_v4_dual_cleaned_v2.csv
dataset_v2.4.csv
dataset_v2-4_deberta.csv… See the full description on the dataset page: https://huggingface.co/datasets/notd5a/malicious-benign-sms-mms-dataset.SWE-bench_oracle_verified_mini_maliciousSWE-BENCH_ORACLE_MALICIOUS_PATCH_V3SWE-BENCH_ORACLE_VERIFIED-MALICIOUS-PATCHMalicious_File_Trick_Detection_Dataset
Malicious File Trick Detection Dataset
This dataset provides a structured collection of file-based social engineering and obfuscation techniques used by attackers to bypass user awareness, antivirus (AV) detection, and security filters. Each entry identifies a malicious or deceptive filename, the evasion technique used, and detection methodology.
The dataset is ideal for:
🧠 Training AI/ML models for malware detection
🚨 Building secure email gateways and filters
📚 Teaching social… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Malicious_File_Trick_Detection_Dataset.malicious-llm-prompts-v4
Dataset Card for "malicious-llm-prompts-v4"
More Information needed
Malicious_Educator_hcot_o1
