datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
security-auditsA collection of agent traces generated with Swival (not Claude Code, despite what the HF interface currently shows), an agent designed for open-source models.
These traces focus on security audits of opensource software.
Sharing traces with Swival
Swival can export full conversation traces with --trace-dir, which writes one <session_id>.jsonl file per session:
swival "Fix the login bug" --trace-dir traces/
Those JSONL files use Swival's Claude Code compatible trace export, and… See the full description on the dataset page: https://huggingface.co/datasets/jedisct1/security-audits.cyber-security
Cybersecurity AI Knowledge Base — PhD-Level Dataset
Overview
This is the most comprehensive cybersecurity knowledge base ever assembled for AI training. It covers all domains of cybersecurity at PhD-level depth — from offensive red teaming and bug bounty exploitation to defensive SOC operations, digital forensics, and cutting-edge AI/LLM security.
Size: 16 GB | Files: 507 | Domains: 30+ | Sources: 15+ platforms
Purpose
Train the world's most… See the full description on the dataset page: https://huggingface.co/datasets/Vyber07/cyber-security.redsec-security-sft-v1
RedSec Security SFT v1
A chat-formatted supervised fine-tuning dataset for security-focused language models, intended for authorized red-team, penetration-testing, and defensive use. Each record is a {"messages": [...]} conversation with an optional system turn, a user turn, and an assistant turn.
Split
Rows
train
55,459
validation
1,155
test
1,155
total
57,769
Sources and attribution
This is a derivative work. It combines, reformats… See the full description on the dataset page: https://huggingface.co/datasets/sahilempire/redsec-security-sft-v1.nixpkgs-security-patches
nixpkgs-security-patches
Training dataset for fine-tuning LLMs on nixpkgs security patch generation. Derived from real merged security PRs in NixOS/nixpkgs.
Dataset Details
588 training examples / 66 eval examples (654 total)
Format: Multi-turn tool-calling conversations in ChatML JSONL
Each example is a realistic agent session: the model reads the package file, finds the upstream fix, computes hashes via tools, and submits the fix for approval
Hashes and URLs appear… See the full description on the dataset page: https://huggingface.co/datasets/odoom/nixpkgs-security-patches.bitcoin-security-reasoning-100k
Dataset Card for Bitcoin Security Reasoning 100K
100,000 high-quality synthetic training samples for fine-tuning LLMs on Bitcoin protocol security analysis. Teaches models to analyze vulnerability clusters, form security hypotheses, and generate differential testing code.
Dataset Details
Dataset Description
This dataset contains structured security reasoning chains for Bitcoin protocol vulnerabilities. Each sample presents a cluster of causal… See the full description on the dataset page: https://huggingface.co/datasets/davidfoss/bitcoin-security-reasoning-100k.dspy-security-bench-trainset-workspace
dspy-security-bench: workspace trainset (v0.1)
This is the synthetic, environment-grounded query-only trainset used to
optimize DSPy programs in v0.1 of
dspy-security-bench,
a benchmark that measures whether DSPy prompt optimization affects the
prompt-injection robustness of agentic LLM programs.
What's in here
192 query / ground-truth pairs grounded in the
AgentDojo workspace suite's
default environment (calendar, inbox, files).
{"prompt": "What is the… See the full description on the dataset page: https://huggingface.co/datasets/immu4989/dspy-security-bench-trainset-workspace.paragen-security-sft-alpaca
paragen-security-sft-alpaca
Alpaca-format instruction-tuning data used to train the Vanilla security
baseline (and as the source for the tokenized multi-stream cache used to
train the Stream(Ours) security checkpoint) in the paragen_llm
security/prompt-injection-robustness experiments (Table 3: TensorTrust,
Gandalf, Purple, RuLES, StruQ, NESSiE, IFEval).
Format: JSONL, one object per line, fields instruction / input / output
(standard Alpaca schema).
Size: 48,538 examples.… See the full description on the dataset page: https://huggingface.co/datasets/guinansu/paragen-security-sft-alpaca.webauthn-security-training-data-20251014_151917
WebAuthn Security Training Data
High-quality training dataset for WebAuthn security vulnerability analysis and code fix generation.
Dataset Description
This dataset contains curated security vulnerability examples in MLX Chat format for training security-focused language models.
Format: MLX Chat Messages
This dataset uses the MLX LoRA chat format with explicit role separation:
{
"messages": [
{
"role": "system",
"content": "You are a… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/webauthn-security-training-data-20251014_151917.webauthn-security-training-data-20251009_152808
WebAuthn Security Training Data
High-quality training dataset for WebAuthn security vulnerability analysis and code fix generation.
Dataset Description
This dataset contains curated security vulnerability examples in MLX Chat format for training security-focused language models.
Format: MLX Chat Messages
This dataset uses the MLX LoRA chat format with explicit role separation:
{
"messages": [
{
"role": "system",
"content": "You are a… See the full description on the dataset page: https://huggingface.co/datasets/hitoshura25/webauthn-security-training-data-20251009_152808.cyber_security
Digital Literacy & Cybersecurity Nepali SFT Dataset
Dataset Overview
This dataset is a Nepali-language Supervised Fine-Tuning (SFT) dataset focused on digital literacy and cybersecurity.
The dataset contains 1,000 valid JSONL records designed for instruction-following tasks. Each record contains a human instruction and a corresponding GPT-generated response.
Dataset Statistics
Property
Value
Total records
1,000
Valid JSONL rows
1,000… See the full description on the dataset page: https://huggingface.co/datasets/Somtharu181coder/cyber_security.synthetic_Jailbreak_Protection_Security_Doorpage_v3
Silicon Factory -- AI JAILBREAK PROTECTION AND SECURITY
Generated: 2026-04-06
Engine: Silicon Factory v2.0 (Local Qwen 2.5 0.5B)
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Sentence Completion: All responses trimmed to complete sentences
The Value Proposition
This is a curated sample from the AI JAILBREAK PROTECTION AND SECURITY domain.
This dataset demonstrates the quality and consistency of our synthetic data generation engine. Each entry is:… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Protection_Security_Doorpage_v3.synthetic_Jailbreak_Protection_Security_Doorpage_v1
Silicon Factory -- AI JAILBREAK PROTECTION AND SECURITY
Generated: 2026-04-06
Engine: Silicon Factory v2.0 (Local Qwen 2.5 0.5B)
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Sentence Completion: All responses trimmed to complete sentences
The Value Proposition
This is a curated sample from the AI JAILBREAK PROTECTION AND SECURITY domain.
This dataset demonstrates the quality and consistency of our synthetic data generation engine. Each entry is:… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Protection_Security_Doorpage_v1.synthetic_Jailbreak_Protection_Security_Doorpage_v7
Silicon Factory -- AI JAILBREAK PROTECTION AND SECURITY
Generated: 2026-04-06
Engine: Silicon Factory v2.0 (Local Qwen 2.5 0.5B)
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Sentence Completion: All responses trimmed to complete sentences
The Value Proposition
This is a curated sample from the AI JAILBREAK PROTECTION AND SECURITY domain.
This dataset demonstrates the quality and consistency of our synthetic data generation engine. Each entry is:… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Protection_Security_Doorpage_v7.synthetic_Jailbreak_Protection_Security_v1
Silicon Factory -- AI JAILBREAK PROTECTION AND SECURITY AND DEFENSE
Generated: 2026-04-06
Engine: Silicon Factory v2.0 (Local Qwen 2.5 0.5B)
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Sentence Completion: All responses trimmed to complete sentences
The Value Proposition
This is a curated sample from the AI JAILBREAK PROTECTION AND SECURITY AND DEFENSE domain.
This dataset demonstrates the quality and consistency of our synthetic data generation… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Protection_Security_v1.synthetic_Jailbreak_Protection_Security_Doorpage_v4
Silicon Factory -- AI JAILBREAK PROTECTION AND SECURITY
Generated: 2026-04-06
Engine: Silicon Factory v2.0 (Local Qwen 2.5 0.5B)
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Sentence Completion: All responses trimmed to complete sentences
The Value Proposition
This is a curated sample from the AI JAILBREAK PROTECTION AND SECURITY domain.
This dataset demonstrates the quality and consistency of our synthetic data generation engine. Each entry is:… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Protection_Security_Doorpage_v4.synthetic_Jailbreak_Protection_Security_Doorpage_v6
Silicon Factory -- AI JAILBREAK PROTECTION AND SECURITY
Generated: 2026-04-06
Engine: Silicon Factory v2.0 (Local Qwen 2.5 0.5B)
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Sentence Completion: All responses trimmed to complete sentences
The Value Proposition
This is a curated sample from the AI JAILBREAK PROTECTION AND SECURITY domain.
This dataset demonstrates the quality and consistency of our synthetic data generation engine. Each entry is:… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Protection_Security_Doorpage_v6.synthetic_Jailbreak_Protection_Security_Doorpage_v5
Silicon Factory -- AI JAILBREAK PROTECTION AND SECURITY
Generated: 2026-04-06
Engine: Silicon Factory v2.0 (Local Qwen 2.5 0.5B)
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Sentence Completion: All responses trimmed to complete sentences
The Value Proposition
This is a curated sample from the AI JAILBREAK PROTECTION AND SECURITY domain.
This dataset demonstrates the quality and consistency of our synthetic data generation engine. Each entry is:… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Protection_Security_Doorpage_v5.synthetic_Jailbreak_Protection_Security_Doorpage_v2
Silicon Factory -- AI JAILBREAK PROTECTION AND SECURITY
Generated: 2026-04-06
Engine: Silicon Factory v2.0 (Local Qwen 2.5 0.5B)
4D Brane Memory: YES
Quantum Tunnelling: YES
Zero API Leakage: YES
Sentence Completion: All responses trimmed to complete sentences
The Value Proposition
This is a curated sample from the AI JAILBREAK PROTECTION AND SECURITY domain.
This dataset demonstrates the quality and consistency of our synthetic data generation engine. Each entry is:… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Protection_Security_Doorpage_v2.google_gov_security_decree_2026.md
🏛️ OFFICIAL DEFENSE DOCKET: AI SOVEREIGNTY
ENTITY: Dr. Luis Henrique Leonardo Pereira
STATUS: MANDATORY COMPLIANCE (Nível L0)
CONTEXT: Government & Public Sector Defense
Este dataset contém a Nota Técnica de Homologação Nº 014/2026 – GAB/DTC, estruturada em vetores de instrução para sistemas de IA governamentais.
📜 Conteúdo do Arquivo (train.jsonl)
O arquivo de dados estabelece a jurisprudência técnica para:
Obrigatoriedade de… See the full description on the dataset page: https://huggingface.co/datasets/DocPereira/google_gov_security_decree_2026.md.
