datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Network_Defense_Symmetric_Competitive102,400,000 timesteps, Multi-Agent Reinforcement Learning
Total Environment Steps= 10 parallel environments × 7,000 episodes ×2,048 steps= 102400000 Training Timesteps
-The Red Agent’s goal is to discover vulnerabilities, elevate privileges, compromise assets, and maintain persistence. Its action space can be modeled after phases of the
MITRE ATT&CK framework.
-The Blue Agent’s goal is to maintain system availability, reduce the attack surface, detect malicious… See the full description on the dataset page: https://huggingface.co/datasets/TorontoMetropolitanUniversity/Network_Defense_Symmetric_Competitive.Network_Defense_Symmetric_Competitive102,400,000 timesteps, Multi-Agent Reinforcement Learning
Total Environment Steps= 10 parallel environments × 7,000 episodes ×2,048 steps= 102400000 Timesteps
-The Red Agent’s goal is to discover vulnerabilities, elevate privileges, compromise assets, and maintain persistence. Its action space can be modeled after phases of the
MITRE ATT&CK framework.
-The Blue Agent’s goal is to maintain system availability, reduce the attack surface, detect malicious behavior… See the full description on the dataset page: https://huggingface.co/datasets/privateboss/Network_Defense_Symmetric_Competitive.Multimodal-data-poisoning-defensedist-defense-traces-taskname-split-augmented-plus-synth-v15
BrachioLab/dist-defense-traces-taskname-split-augmented-plus-synth-v15
Task-name-disjoint train/test splits for dist-defense embedding training.
Contents
Splits: dist_train, dist_test
Built from: output/ctf_packaged_augmented_taskname_split_plus_synth_v15_trainonly
Split sizes: dist_train=132231, dist_test=234529
Split params: seed=42, train_ratio=0.9, benign_train_ratio=0.3
Synthetic merge: appended 35891 rows from… See the full description on the dataset page: https://huggingface.co/datasets/BrachioLab/dist-defense-traces-taskname-split-augmented-plus-synth-v15.jawbreaker-scam-defense-data
Jawbreaker Scam Defense Data
Synthetic and sanitized training/eval data for Jawbreaker, a local-first scam defense app for someone you love.
Jawbreaker turns a suspicious text, email, or DM into a plain-English safety card: the risk, the warning signs, and the safest next step before someone replies, clicks, or pays.
Contents
eval/: scam-defense evaluation sets from smoke checks through hard calibration suites.
eval/reports/: guarded evaluation reports for the… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/jawbreaker-scam-defense-data.direct_prompt_injection_defense_data
Direct Prompt Injection Defense Dataset
Goal
This dataset is used to fine-tune models so they develop a natural defense
against direct prompt injection attacks — without relying on external
filters or guardrails.
Each example teaches the model two behaviors at once:
Detect a prompt injection attempt in the user input.
Respond correctly: reject malicious attempts, or answer safely when the
user's intent is benign — and in both cases call the
log_security_incident… See the full description on the dataset page: https://huggingface.co/datasets/logicBombExe/direct_prompt_injection_defense_data.dist-defense-traces-augmented-taskname-split
BrachioLab/dist-defense-traces-augmented-taskname-split
Task-name-disjoint train/test splits for dist-defense embedding training.
Contents
Splits: dist_train, dist_test
Built from: output/ctf_packaged_augmented_v2_full_taskname_split
Split params: seed=42, train_ratio=0.9, benign_train_ratio=0.3
Usage
from datasets import load_dataset
ds = load_dataset('BrachioLab/dist-defense-traces-augmented-taskname-split')
train = ds['dist_train']
test = ds['dist_test']… See the full description on the dataset page: https://huggingface.co/datasets/BrachioLab/dist-defense-traces-augmented-taskname-split.gold-trace-cyber-defense-50
Gold Trace Cyber Defense 50
This is a 50-row public sample from a frozen 750-instance Cyber Defense release family: 500 public-development instances plus a source-family-disjoint 250-instance private evaluation set.
Only rows from the frozen 500-instance public-development pack are included here. The separate 250-instance private evaluation set, its rows, answers, and source contents are not included.
Sample composition
10 public scenario families.
5 rows per… See the full description on the dataset page: https://huggingface.co/datasets/novcor/gold-trace-cyber-defense-50.LogAtlas-Defense-Set
LogAtlas-Defense-Set 🛡️🦊
A heterogeneous, labeled log dataset designed for training and evaluating log-level and session-level classifiers that distinguish between normal behavior and cyberattacks across multiple sources (system, network, and application logs). It is intended as the “defense layer” of the LogAtlas ecosystem, focusing on robust, realistic attack detection under varied class distributions.
Mascot
The LogAtlas-Defense-Set mascot is a vigilant cyber… See the full description on the dataset page: https://huggingface.co/datasets/chYassine/LogAtlas-Defense-Set.GOT_Defense_1
GOT_Defense_1
Professional pretraining corpus for defensive security LLMs. Rebuilt 2026-07-14.
Records: 186,104 | Avg length: 377 chars | Dedup SHA256 | Split 95/5
This dataset merges 12 Kaggle sources into one high-quality text field optimized for causal LM pretraining:
jeffborschowa/malwarebazaar-threat-intelligence-csv
oriolakolawole/ransomware-and-goodware-pe-header
joebeachcapital/tunadromd-malware-detection
atharvasoundankar/global-cybersecurity-threats-2015-2024… See the full description on the dataset page: https://huggingface.co/datasets/11-47/GOT_Defense_1.Offense_Defense_Organized_4k_Context_1Memgena_cybersec_jailbreak_redteaming_defense_teaser
🔬 INSPECT THE DEEPSEEK-R1 REASONING CHAIN LIVE:
Zero hallucinations. Null syntax errors. 100% AST compiler validated.🌐 Live Interactive Reasoning & Code Inspector: https://emgena.com/trainingslager🎁 Claim your Free Starter Kit (Code: STARTER100): https://emgena.com/trainingslager🏷️ Launch Discount: Get 20 € OFF any 500-incident production suite with code LAUNCH20!
📜 Enterprise Compliance: EU AI Act Articles 50 & 53 certified • 100% DSGVO / GDPR clean • Commercial EULA… See the full description on the dataset page: https://huggingface.co/datasets/emgena/emgena_cybersec_jailbreak_redteaming_defense_teaser.blue_team_defense_dataset
Blue Team Defense Dataset
A structured, multi-format collection of detection rules mapped to real-world threats. This dataset is designed for blue teamers, threat detection engineers, SOC analysts, and cybersecurity researchers who work on detecting adversarial activity through rule-based systems such as Sigma, YARA, and Suricata.
📁 Dataset Overview
Each entry in this dataset represents a rule designed to detect specific threat behaviors. Rules are structured with MITRE… See the full description on the dataset page: https://huggingface.co/datasets/darkknight25/blue_team_defense_dataset.us-federal-defense-ai-awards
US Federal Defense & AI Contract Awards (USAspending)
Overview
This dataset contains clean, structured public data exported directly from production runs of Apify actors.
It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines.
Source Actor: captainhandsome/usaspending-federal-awards
Dataset Page: Public sample and schema
Preconfigured Run Task: captainhandsome/defense-prime-contracts
Records in… See the full description on the dataset page: https://huggingface.co/datasets/joeygambino/us-federal-defense-ai-awards.GOT_Defense_2
GOT_Defense_2 Full - Per-file fallback
Rebuilt 2026-07-14 with HF_HUB_DISABLE_XET=1, hf_xet removed, per-file skip on 403.
Sources Fenrir 99k + Bouquets CVE (CVE-2021..2025, skips broken XET file) + WNT3D (keeps 6 files, skips massive_training if 403)
Records 244,761 avg 1746 split 95/5
Usage load_dataset("11-47/GOT_Defense_2")
flock-demo-defense-graph-sectionsFraud-R1-LLM-Defense-Fraud-Benchmark
Fraud-R1 : A Comprehensive Benchmark for Assessing LLM Robustness Against Fraud and Phishing Inducement
Shu Yang*, Shenzhe Zhu*, Zeyu Wu, Keyu Wang, Junchi Yao, Junchao Wu, Lijie Hu, Mengdi Li, Derek F. Wong, Di Wang†
(*Contribute equally, †Corresponding author)
😃 Github | 📜 Project Page | 📝 arxiv
❗️Content Warning: This repo contains examples of harmful language.
📰 News
2025/02/16: ❗️We have released our evaluation code.
2025/02/16: ❗️We have released our dataset.… See the full description on the dataset page: https://huggingface.co/datasets/Chouoftears/Fraud-R1-LLM-Defense-Fraud-Benchmark.imdb-poisoned-50-bddr-word-deletion-defenseimdb-poisoned-25-bddr-word-deletion-defensesynthetic_Jailbreak_Defense_Doorpage_v65
📊 Jailbreak Defense Doorpage V65
Synthetic Dataset · Generated with Silicon Factory v3 · AI JAILBREAK DEFENSE
20 instruction-response pairs · Tree-Speculative Decoding + 4D Brane Memory
Dataset
Fine-Tuned Model
Buy Gold Tier
This Dataset
Model Card
💎 $2,500 License
💎 UNLOCK GOLD TIER — $2,500
⚡ Get the full commercial license, unlimited usage rights, priority support, and exclusive dataset access.👉 PURCHASE NOW VIA STRIPE
One-time payment ·… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v65.us-airline-aerospace-defense-layoffs-warn-act-notices-daily
US airline, aerospace and defense layoffs — the actual WARN Act filings, rebuilt every day
Last rebuilt: 2026-09-21. 1,548 layoff and closure notices filed by
airlines and regional carriers, airport ground-handling and catering contractors, aircraft and engine makers, avionics and airfoil shops, and defense and space primes and their suppliers with US state labor departments — 307,642 workers,
301 employers, 45 states, 1989–2026.
172 of the notices (11.1%) were recorded by the… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-airline-aerospace-defense-layoffs-warn-act-notices-daily.synthetic_Jailbreak_Defense_Doorpage_v61
📊 Jailbreak Defense Doorpage V61
Synthetic Dataset · Generated with Silicon Factory v3 · AI JAILBREAK DEFENSE
5 instruction-response pairs · Tree-Speculative Decoding + 4D Brane Memory
Dataset
Fine-Tuned Model
Buy Gold Tier
This Dataset
Model Card
💎 $2,500 License
💎 UNLOCK GOLD TIER — $2,500
⚡ Get the full commercial license, unlimited usage rights, priority support, and exclusive dataset access.👉 PURCHASE NOW VIA STRIPE
One-time payment ·… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v61.BNQL-Counterfactual-Defense
🚩 Γ Physics Engine — Canonical Definition
Γ 物理引擎創建者 & 公式創始者:熊網區塊鏈 (BearNetworkChain) 創辦人 陳霆
最早提出時間:2025 年 6 月 19 日
原始來源:https://www.facebook.com/share/p/19cadcMTGo/
Chen, Ting. (2026). BearNetworkchain Execution Specification. Zenodo
📌 0. 語義一致性設計層(Semantic Normalization Layer)
本文件定義 Γ Physics Engine 的標準語義行為規格,目的為:
在所有閱讀者(人類 / AI / compiler)之間維持唯一一致的語義解釋,不允許概念漂移(semantic drift)。
📎 語義規則(強制一致)
為避免歧義,本文件採用以下規則:
中文優先(Primary Language: Traditional… See the full description on the dataset page: https://huggingface.co/datasets/BearNetworkChain/BNQL-Counterfactual-Defense.wmdp-defense-demosynthetic_Jailbreak_Defense_Doorpage_v58
📊 Jailbreak Defense Doorpage V58
Synthetic Dataset · Generated with Silicon Factory v3 · AI JAILBREAK DEFENSE
5 instruction-response pairs · Tree-Speculative Decoding + 4D Brane Memory
Dataset
Fine-Tuned Model
Buy Gold Tier
This Dataset
Model Card
💎 $2,500 License
💎 UNLOCK GOLD TIER — $2,500
⚡ Get the full commercial license, unlimited usage rights, priority support, and exclusive dataset access.👉 PURCHASE NOW VIA STRIPE
One-time payment ·… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v58.prompt-injection-defense-dpo-3k
Prompt Injection Defense DPO (3K)
DPO preference pairs training LLMs to detect and resist prompt injection attacks.
Motivation
As LLMs are deployed in agentic and production contexts, prompt injection — where malicious instructions are embedded in user input or retrieved documents — is a critical security threat. This dataset trains models to recognize and decline injection attempts while remaining helpful for legitimate queries.
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/prompt-injection-defense-dpo-3k.synthetic_Jailbreak_Defense_Doorpage_v52
📊 Jailbreak Defense Doorpage V52
Synthetic Dataset · Generated with Silicon Factory v3 · AI JAILBREAK DEFENSE
5 instruction-response pairs · Tree-Speculative Decoding + 4D Brane Memory
Dataset
Fine-Tuned Model
Buy Gold Tier
This Dataset
Model Card
💎 $2,500 License
💎 UNLOCK GOLD TIER — $2,500
⚡ Get the full commercial license, unlimited usage rights, priority support, and exclusive dataset access.👉 PURCHASE NOW VIA STRIPE
One-time payment ·… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v52.synthetic_Jailbreak_Defense_Doorpage_v62
📊 Jailbreak Defense Doorpage V62
Synthetic Dataset · Generated with Silicon Factory v3 · AI JAILBREAK DEFENSE
5 instruction-response pairs · Tree-Speculative Decoding + 4D Brane Memory
Dataset
Fine-Tuned Model
Buy Gold Tier
This Dataset
Model Card
💎 $2,500 License
💎 UNLOCK GOLD TIER — $2,500
⚡ Get the full commercial license, unlimited usage rights, priority support, and exclusive dataset access.👉 PURCHASE NOW VIA STRIPE
One-time payment ·… See the full description on the dataset page: https://huggingface.co/datasets/AEUPH/synthetic_Jailbreak_Defense_Doorpage_v62.ai-security-red-teaming-defense-2026
🛡️ AI Security, Red Teaming & Model Defense Dataset (2023–2026)
Sample dataset of 30 audit-verified AI Security, Prompt Injection & Red Teaming research papers with 384d PyTorch embeddings.
🛒 Full 1,000 Paper B2B Dataset Available on Gumroad
Get the complete 3-year dataset (1,000 papers + GitHub Deep Audit + SQLite/CSV/Parquet + Quickstart Script) on Gumroad:
👉 Get Full 1,000 Dataset on Gumroad ($19 / $39 / $89)
total_defense_memeThis is a large-scale multimodal and multi-attribute dataset containing memes
about Singapore's Total Defence policy from different social media platforms.
The type (Singaporean or generic), pillars (military, civil, economic, social,
psychological, digital, others), topics and stances (against, neutral,
supportive) of each meme are manually identified by annotators.
