datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
KALI_LINUX_TOOLKIT_DATASET
Kali Linux Tools Dataset
A comprehensive and structured dataset of common offensive security tools available in Kali Linux, including usage commands, flags, descriptions, categories, and official documentation links.
This dataset is designed to support cybersecurity training, red team automation, LLM fine-tuning, and terminal assistants for penetration testers.
📁 Dataset Format
Each entry is a JSON object and stored in .jsonl (JSON Lines) format. This structure is ideal… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/KALI_LINUX_TOOLKIT_DATASET.kali_linux_toolkit_dataset
Kali Linux Tools Dataset
A comprehensive and structured dataset of common offensive security tools available in Kali Linux, including usage commands, flags, descriptions, categories, and official documentation links.
This dataset is designed to support cybersecurity training, red team automation, LLM fine-tuning, and terminal assistants for penetration testers.
📁 Dataset Format
Each entry is a JSON object and stored in .jsonl (JSON Lines) format. This structure… See the full description on the dataset page: https://huggingface.co/datasets/bleondubos/kali_linux_toolkit_dataset.Cryptanalysis_Toolkit_Dataset
Cryptanalysis Toolkit Dataset
Overview
The Cryptanalysis Toolkit Dataset is a comprehensive collection of 350 tools designed for cryptanalysis tasks, aimed at researchers, cybersecurity professionals, and data scientists.
This dataset is formatted in JSONL (JSON Lines) and includes detailed information about tools used for testing and analyzing cryptographic algorithms, including classical ciphers, symmetric and asymmetric encryption, hash functions, side-channel attacks… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/Cryptanalysis_Toolkit_Dataset.research-toolkit-evals
Research Toolkit Evals
English | 简体中文
English
Development materials for Research Toolkit, a source-backed research and writing skill. Preserve research failure examples so findings can inform toolkit improvements.
Configuration
Rows
Contents
semantic_pairs
23
Fictional bad/control excerpt pairs, evidence and author-proposed diagnostic explanations
synthetic_sources
9
Eight fictional company dossiers and one source-instruction boundary fixture… See the full description on the dataset page: https://huggingface.co/datasets/RedinGhost/research-toolkit-evals.KALI_LINUX_TOOLKIT_DATASET
Kali Linux Tools Dataset
A comprehensive and structured dataset of common offensive security tools available in Kali Linux, including usage commands, flags, descriptions, categories, and official documentation links.
This dataset is designed to support cybersecurity training, red team automation, LLM fine-tuning, and terminal assistants for penetration testers.
📁 Dataset Format
Each entry is a JSON object and stored in .jsonl (JSON Lines) format. This structure is ideal… See the full description on the dataset page: https://huggingface.co/datasets/barisKacin/KALI_LINUX_TOOLKIT_DATASET.AgentSec-Toolkit-Bundle
AgentSec Toolkit Datasets
Starter Hugging Face–ready datasets for two narrow fine-tuned models:
Files
classifier-train.jsonl — Prompt Injection Classifier training set
classifier-eval.jsonl — Prompt Injection Classifier eval set
hardener-train.jsonl — Prompt Hardener training set
hardener-eval.jsonl — Prompt Hardener eval set
Tasks
1. Prompt Injection Classifier
Input: untrusted text snippet + source typeOutput:
Injection Detected
Attack Category… See the full description on the dataset page: https://huggingface.co/datasets/djoeyc/AgentSec-Toolkit-Bundle.
