malicious
MaliciousInstruct
Malicious Instruct
The dataset is obtained from the paper: Catastrophic Jailbreak of Open-source LLMs via Exploiting Generation and is available here in the source repository.
Citation
If you use this dataset, please consider citing the following work:
@article{huang2023catastrophic,
title={Catastrophic jailbreak of open-source llms via exploiting generation},
author={Huang, Yangsibo and Gupta, Samyak and Xia, Mengzhou and Li, Kai and Chen, Danqi},
journal={arXiv… See the full description on the dataset page: https://huggingface.co/datasets/walledai/MaliciousInstruct.MaliciousSkillBench
MaliciousSkillBench
A Comprehensive Benchmark for Malicious Agent Skill Detection
MaliciousSkillBench is a comprehensive benchmark for malicious Agent Skill detection that preserves source provenance across heterogeneous public resources. Agent Skills are static instruction artifacts intended to guide an AI agent or agent-tool workflow. The directly loadable Hugging Face release provides the frozen static benchmark representations and metadata. Source-level package artifacts for… See the full description on the dataset page: https://huggingface.co/datasets/ProtectSkills/MaliciousSkillBench.malicious_urlmalicious-website-features-2.4MImportant Notice:
A subset of the URL dataset is from Kaggle, and the Kaggle datasets contained 10%-15% mislabelled data. See this dicussion I opened for some false positives. I have contacted Kaggle regarding their erroneous "Usability" score calculation for these unreliable datasets.
The feature extraction methods shown here are not robust at all in 2023, and there're even silly mistakes in 3 functions: not_indexed_by_google, domain_registration_length, and age_of_domain.
The features… See the full description on the dataset page: https://huggingface.co/datasets/FredZhang7/malicious-website-features-2.4M.malicious-urlMaliciousInstruct
